- Hardware
- 4× NVIDIA GeForce RTX 5090
- Workload
- 1024 → 256
- Runs
- 8 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=64 | 2,327.4 | 2,528.4 ms | 25.04 ms | 7,674.9 ms | 0 | run_lPHnsM7weInxjd_C | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=16 | 1,329.5 | 652.7 ms | 11.48 ms | 3,069.3 ms | 0 | run_QjI5OdTC3LlWDCBs | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=4 | 617.7 | 194 ms | 6.26 ms | 1,662.1 ms | 0 | run_8FxfHM5XZUC0ZUa6 | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=2 | 323.3 | 109.3 ms | 5.98 ms | 1,585.2 ms | 0 | run_QSDg-z_HDZUQlhfi | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=128 | 3,018.5 | 4,129.3 ms | 40.97 ms | 13,044.9 ms | 0 | run_gfWaR-EN7-iG8d8f | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=32 | 1,277.3 | 1,269 ms | 59.17 ms | 15,730.2 ms | 0 | run_F_ID-98IPqyMT3Zc | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=8 | 944.6 | 343.9 ms | 8.25 ms | 2,179.6 ms | 0 | run_ExfD7Ne8gmqWHZBf | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4 | c=1 | 235 | 58.1 ms | 4.05 ms | 1,090 ms | 0 | run_0fdYMWzrl34_j431 |
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=64
- Req/s
- Output tok/s
- 2,327.4
- p95 TTFT / TPOT
- 2,528.4 ms / 25.04 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=16
- Req/s
- Output tok/s
- 1,329.5
- p95 TTFT / TPOT
- 652.7 ms / 11.48 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=4
- Req/s
- Output tok/s
- 617.7
- p95 TTFT / TPOT
- 194 ms / 6.26 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=2
- Req/s
- Output tok/s
- 323.3
- p95 TTFT / TPOT
- 109.3 ms / 5.98 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=128
- Req/s
- Output tok/s
- 3,018.5
- p95 TTFT / TPOT
- 4,129.3 ms / 40.97 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=32
- Req/s
- Output tok/s
- 1,277.3
- p95 TTFT / TPOT
- 1,269 ms / 59.17 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=8
- Req/s
- Output tok/s
- 944.6
- p95 TTFT / TPOT
- 343.9 ms / 8.25 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4- Test load
- c=1
- Req/s
- Output tok/s
- 235
- p95 TTFT / TPOT
- 58.1 ms / 4.05 ms