- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 8192 → 256
- Runs
- 9 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=4 | 305.6 | 2,484.1 ms | 10.89 ms | — | 0 | run_o9RGP_25BQeQR59o | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=3 | 304 | 1,302 ms | 13.43 ms | — | 0 | run_3BDbEa4p7fG_ces3 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 274.2 | 1,186.1 ms | 7.78 ms | — | 0 | run_baQvEtRupq-RH4yj | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 207.8 | 648.6 ms | 3.94 ms | — | 0 | run_j7-68fHeePgq1RaT | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 207.9 | 647.5 ms | 3.94 ms | — | 0 | run_RP8_yJMtqc-rJRKu | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 208 | 649.9 ms | 3.93 ms | — | 0 | run_wI6gPdjut78iyVVR | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 208.1 | 648.6 ms | 3.93 ms | — | 0 | run_qtIM7sYO5VUQ2HmX | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 208.1 | 648.7 ms | 3.93 ms | — | 0 | run_M7cqO7Ecb-elWHRy | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 207.9 | 651.4 ms | 3.94 ms | — | 0 | run_Af-H_BUeFX_g2m6m |
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=4
- Req/s
- Output tok/s
- 305.6
- p95 TTFT / TPOT
- 2,484.1 ms / 10.89 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=3
- Req/s
- Output tok/s
- 304
- p95 TTFT / TPOT
- 1,302 ms / 13.43 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 274.2
- p95 TTFT / TPOT
- 1,186.1 ms / 7.78 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 207.8
- p95 TTFT / TPOT
- 648.6 ms / 3.94 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 207.9
- p95 TTFT / TPOT
- 647.5 ms / 3.94 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 208
- p95 TTFT / TPOT
- 649.9 ms / 3.93 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 208.1
- p95 TTFT / TPOT
- 648.6 ms / 3.93 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 208.1
- p95 TTFT / TPOT
- 648.7 ms / 3.93 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 207.9
- p95 TTFT / TPOT
- 651.4 ms / 3.94 ms