- Hardware
- 1× NVIDIA GeForce RTX 5090
- Draft model
- incoai/Qwen3.8-27B-DFlash2
- Workload
- 8192 → 256
- Runs
- 9 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=4 | 296.9 | 2,140.2 ms | 11.14 ms | — | 0 | run_G0VEc4dkOcJdfmOf | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=3 | 299.6 | 1,278.1 ms | 11.28 ms | — | 0 | run_hL-HgU1CB7sW17kN | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 267 | 675.3 ms | 8.14 ms | — | 0 | run_IvyXtniXDLmTUbSp | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 267.1 | 674.6 ms | 8.13 ms | — | 0 | run_U6iGiktsmhwdljNA | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 266.7 | 686.7 ms | 8.16 ms | — | 0 | run_PMnqSjjH5rN-YNSn | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=1 | 202.4 | 651.2 ms | 3.46 ms | — | 0 | run_QaFexjkwufHT-MQt | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 203.5 | 649.9 ms | 3.42 ms | — | 0 | run_Tx838CoIbYscQaJv | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 203.5 | 651.2 ms | 3.42 ms | — | 0 | run_5kGX-v6JSVmHFjfW | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 203.5 | 650.6 ms | 3.42 ms | — | 0 | run_XS6QTT-rtZmhQye3 |
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=4
- Req/s
- Output tok/s
- 296.9
- p95 TTFT / TPOT
- 2,140.2 ms / 11.14 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=3
- Req/s
- Output tok/s
- 299.6
- p95 TTFT / TPOT
- 1,278.1 ms / 11.28 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 267
- p95 TTFT / TPOT
- 675.3 ms / 8.14 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 267.1
- p95 TTFT / TPOT
- 674.6 ms / 8.13 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 266.7
- p95 TTFT / TPOT
- 686.7 ms / 8.16 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=1
- Req/s
- Output tok/s
- 202.4
- p95 TTFT / TPOT
- 651.2 ms / 3.46 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 203.5
- p95 TTFT / TPOT
- 649.9 ms / 3.42 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 203.5
- p95 TTFT / TPOT
- 651.2 ms / 3.42 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 203.5
- p95 TTFT / TPOT
- 650.6 ms / 3.42 ms