- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 8192 → 256
- Runs
- 9 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=4 | 180.2 | 2,351.7 ms | 19.78 ms | — | 0 | run_yK-07d4Y5vGe87ys | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=3 | 151.6 | 1,769 ms | 17.25 ms | — | 0 | run_XSLaVc1L7v1wr5aZ | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 117.1 | 1,186.6 ms | 14.61 ms | — | 0 | run_kG55hY85vGbhzWhN | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 73.1 | 610.2 ms | 11.38 ms | — | 0 | run_Y24WiWY3H_uECWVA | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 73.2 | 610.8 ms | 11.38 ms | — | 0 | run_z8t4rhaSqIWGcZof | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 73 | 610.5 ms | 11.4 ms | — | 0 | run_YD-0nGHOrlglf8qf | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 73.3 | 610.2 ms | 11.37 ms | — | 0 | run_PT98anCagtktLf9Y | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 73.3 | 610.1 ms | 11.37 ms | — | 0 | run_zfUui30GNOyRV1S6 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 73.1 | 610.4 ms | 11.4 ms | — | 0 | run_ACIJykREf1YrdTgz |
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=4
- Req/s
- Output tok/s
- 180.2
- p95 TTFT / TPOT
- 2,351.7 ms / 19.78 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=3
- Req/s
- Output tok/s
- 151.6
- p95 TTFT / TPOT
- 1,769 ms / 17.25 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 117.1
- p95 TTFT / TPOT
- 1,186.6 ms / 14.61 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 73.1
- p95 TTFT / TPOT
- 610.2 ms / 11.38 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 73.2
- p95 TTFT / TPOT
- 610.8 ms / 11.38 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 73
- p95 TTFT / TPOT
- 610.5 ms / 11.4 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 73.3
- p95 TTFT / TPOT
- 610.2 ms / 11.37 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 73.3
- p95 TTFT / TPOT
- 610.1 ms / 11.37 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 73.1
- p95 TTFT / TPOT
- 610.4 ms / 11.4 ms