- Hardware
- 4× NVIDIA GeForce RTX 5090
- Workload
- 1024 → 256
- Runs
- 8 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=64 | 6,513.6 | 617.3 ms | 10.63 ms | 2,931.1 ms | 0 | run_mfqbwBkSDDFoKOa3 | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=16 | 3,231.6 | 204 ms | 5.12 ms | 1,420.6 ms | 0 | run_qWxZVLWWpoRHNNAV | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=4 | 1,498.5 | 53.5 ms | 2.48 ms | 685.5 ms | 0 | run_3rdXHqlbZggWv4hH | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=2 | 750.1 | 53.5 ms | 2.48 ms | 684.4 ms | 0 | run_hDHgO-XcKjZB9yaq | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=128 | 7,420.9 | 5,167 ms | 9.73 ms | 6,802.6 ms | 0 | run_DapqPLH83TG71M1Y | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=32 | 4,385.9 | 329.7 ms | 7.69 ms | 2,234.2 ms | 0 | run_KyMqf9qjbzdX1J7S | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=8 | 2,150.7 | 109.8 ms | 3.7 ms | 1,120.8 ms | 0 | run_bOQg5EwiyS1jpvmn | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=1 | 375.4 | 53 ms | 2.48 ms | 684.4 ms | 0 | run_mHbNqg449CsCxEIt |
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=64
- Req/s
- Output tok/s
- 6,513.6
- p95 TTFT / TPOT
- 617.3 ms / 10.63 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=16
- Req/s
- Output tok/s
- 3,231.6
- p95 TTFT / TPOT
- 204 ms / 5.12 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=4
- Req/s
- Output tok/s
- 1,498.5
- p95 TTFT / TPOT
- 53.5 ms / 2.48 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=2
- Req/s
- Output tok/s
- 750.1
- p95 TTFT / TPOT
- 53.5 ms / 2.48 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=128
- Req/s
- Output tok/s
- 7,420.9
- p95 TTFT / TPOT
- 5,167 ms / 9.73 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=32
- Req/s
- Output tok/s
- 4,385.9
- p95 TTFT / TPOT
- 329.7 ms / 7.69 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=8
- Req/s
- Output tok/s
- 2,150.7
- p95 TTFT / TPOT
- 109.8 ms / 3.7 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=1
- Req/s
- Output tok/s
- 375.4
- p95 TTFT / TPOT
- 53 ms / 2.48 ms