- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 256 → 1024
- Runs
- 10 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 656.6 | 139.8 ms | 12.07 ms | — | 0 | run_FYzLUPDhk2Yc4ZFS | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 656.3 | 140.1 ms | 12.08 ms | — | 0 | run_qT-6TMkmPwMdxd6z | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 656.4 | 139.9 ms | 12.08 ms | — | 0 | run_BzUJKLOpT1foCQmQ | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=6 | 487 | 109.9 ms | 11.85 ms | — | 0 | run_OcdPjiCmmOszJW6_ | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=4 | 319.5 | 76 ms | 12.47 ms | — | 0 | run_PyERufoSv0dDFq0g | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 164.1 | 48.5 ms | 12.17 ms | — | 0 | run_PcbLolgZAS2mVFJt | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 88.9 | 38.1 ms | 11.22 ms | — | 0 | run_jjgZ3n3VKgVVSkgA | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 89.4 | 35.8 ms | 11.19 ms | — | 0 | run_riEWQzTB5E6iHuP9 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 89.4 | 37.7 ms | 11.19 ms | — | 0 | run_s05fh5tQpBVA9sI7 | |
| sglang 0.0.0.dev1+g5f55db35e | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 89.2 | 37.8 ms | 11.22 ms | — | 0 | run_rmh6e3v_PZG5BjWM |
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 656.6
- p95 TTFT / TPOT
- 139.8 ms / 12.07 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 656.3
- p95 TTFT / TPOT
- 140.1 ms / 12.08 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 656.4
- p95 TTFT / TPOT
- 139.9 ms / 12.08 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=6
- Req/s
- Output tok/s
- 487
- p95 TTFT / TPOT
- 109.9 ms / 11.85 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=4
- Req/s
- Output tok/s
- 319.5
- p95 TTFT / TPOT
- 76 ms / 12.47 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 164.1
- p95 TTFT / TPOT
- 48.5 ms / 12.17 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 88.9
- p95 TTFT / TPOT
- 38.1 ms / 11.22 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 89.4
- p95 TTFT / TPOT
- 35.8 ms / 11.19 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 89.4
- p95 TTFT / TPOT
- 37.7 ms / 11.19 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 89.2
- p95 TTFT / TPOT
- 37.8 ms / 11.22 ms