- Hardware
- 1× NVIDIA GeForce RTX 5090
- Draft model
- incoai/Qwen3.8-27B-DFlash2
- Workload
- 4096 → 1
- Runs
- 7 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=8 | 3.5 | 2,356.6 ms | 0 ms | — | 0 | run_Qap9Q74gQK9MNqDa | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=6 | 3.5 | 1,726.1 ms | 0 ms | — | 0 | run_rmav0sFoNhpjzv8l | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=4 | 3.5 | 1,208.7 ms | 0 ms | — | 0 | run_e2Y0q3TGsJQYEFGD | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 3.5 | 629.7 ms | 0 ms | — | 0 | run_u3pyRlo_XXZa_BeW | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 3.5 | 629.5 ms | 0 ms | — | 0 | run_u6fAm0QtXs6oqOy- | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 3.5 | 631 ms | 0 ms | — | 0 | run_APqikSrVvuhXXHp1 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=1 | 3.4 | 295.3 ms | 0 ms | — | 0 | run_vtQdD8fySOgDVZS5 |
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=8
- Req/s
- Output tok/s
- 3.5
- p95 TTFT / TPOT
- 2,356.6 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=6
- Req/s
- Output tok/s
- 3.5
- p95 TTFT / TPOT
- 1,726.1 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=4
- Req/s
- Output tok/s
- 3.5
- p95 TTFT / TPOT
- 1,208.7 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 3.5
- p95 TTFT / TPOT
- 629.7 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 3.5
- p95 TTFT / TPOT
- 629.5 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 3.5
- p95 TTFT / TPOT
- 631 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=1
- Req/s
- Output tok/s
- 3.4
- p95 TTFT / TPOT
- 295.3 ms / 0 ms