- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 4096 → 1
- Runs
- 7 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 3.7 | 2,209.9 ms | 0 ms | — | 0 | run_FxKGqjdkjx14BwR- | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=6 | 3.7 | 1,668.4 ms | 0 ms | — | 0 | run_JMb-m604g1F6rFj6 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=4 | 3.7 | 1,131.7 ms | 0 ms | — | 0 | run_aMD1tPYsFLz_1wVn | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 3.7 | 595.7 ms | 0 ms | — | 0 | run_PxlhuqS2TEpaqDA1 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 3.7 | 595 ms | 0 ms | — | 0 | run_OJCpIlKXcUAitTF9 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 3.7 | 596.2 ms | 0 ms | — | 0 | run_xitx0BeMs3KHD2V7 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 3.6 | 275.7 ms | 0 ms | — | 0 | run_8CkbuHSlqfBM8f0X |
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 3.7
- p95 TTFT / TPOT
- 2,209.9 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=6
- Req/s
- Output tok/s
- 3.7
- p95 TTFT / TPOT
- 1,668.4 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=4
- Req/s
- Output tok/s
- 3.7
- p95 TTFT / TPOT
- 1,131.7 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 3.7
- p95 TTFT / TPOT
- 595.7 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 3.7
- p95 TTFT / TPOT
- 595 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 3.7
- p95 TTFT / TPOT
- 596.2 ms / 0 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 3.6
- p95 TTFT / TPOT
- 275.7 ms / 0 ms