- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 256 → 1024
- Runs
- 10 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 2,329.1 | 1,235.8 ms | 5.24 ms | — | 0 | run_ucnXL4xlzYB0WExx | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=6 | 1,986.2 | 117.7 ms | 6.08 ms | — | 0 | run_JEq9OLfaQCO9CHzn | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=6 | 1,986.7 | 117.4 ms | 6.08 ms | — | 0 | run_l27XVarmQiKyGJf9 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=6 | 1,914.7 | 109.8 ms | 4.44 ms | — | 0 | run_b9VV0WgkqDwEwZ1z | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=4 | 1,527 | 91.1 ms | 3.18 ms | — | 0 | run_ZZ4xMERgyAoHQgGr | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 854.3 | 58.1 ms | 4.06 ms | — | 0 | run_Y8t06MkjhQ5Hbobt | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 436.7 | 43.6 ms | 2.84 ms | — | 0 | run_pjgVwVQ4fyoDIT9Y | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 436.7 | 43.6 ms | 2.84 ms | — | 0 | run_Y055cveCFjmgKavh | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 436.7 | 43.6 ms | 2.84 ms | — | 0 | run_pVFDaY3ZKHpMkDRJ | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 435.8 | 43.7 ms | 2.85 ms | — | 0 | run_5zWNumbaU3hXyqUw |
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 2,329.1
- p95 TTFT / TPOT
- 1,235.8 ms / 5.24 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=6
- Req/s
- Output tok/s
- 1,986.2
- p95 TTFT / TPOT
- 117.7 ms / 6.08 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=6
- Req/s
- Output tok/s
- 1,986.7
- p95 TTFT / TPOT
- 117.4 ms / 6.08 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=6
- Req/s
- Output tok/s
- 1,914.7
- p95 TTFT / TPOT
- 109.8 ms / 4.44 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=4
- Req/s
- Output tok/s
- 1,527
- p95 TTFT / TPOT
- 91.1 ms / 3.18 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 854.3
- p95 TTFT / TPOT
- 58.1 ms / 4.06 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 436.7
- p95 TTFT / TPOT
- 43.6 ms / 2.84 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 436.7
- p95 TTFT / TPOT
- 43.6 ms / 2.84 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 436.7
- p95 TTFT / TPOT
- 43.6 ms / 2.84 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 435.8
- p95 TTFT / TPOT
- 43.7 ms / 2.85 ms