- Hardware
- 4× NVIDIA GeForce RTX 5090
- Workload
- 1024 → 256
- Runs
- 8 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=64 | 6,157.7 | 427.9 ms | 10.55 ms | 2,895.7 ms | 0 | run_vR-aN6dQN7oN9vEX | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=16 | 2,082.9 | 121.9 ms | 7.53 ms | 1,977.9 ms | 0 | run_aFn0Jx0E_Ml6eRgq | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=4 | 725.3 | 42 ms | 5.39 ms | 1,415.4 ms | 0 | run_TLgC2HLVcevQOitz | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=2 | 362.6 | 41.5 ms | 5.38 ms | 1,413.2 ms | 0 | run_2Hz338t5HaGnvCAY | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=128 | 7,982.7 | 3,587.3 ms | 13.63 ms | 6,850.7 ms | 0 | run_H2PYudiZpc26_sUQ | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=32 | 3,570.5 | 219.7 ms | 8.58 ms | 2,282 ms | 0 | run_ELXdCg4-RIo_ojeG | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=8 | 1,209.9 | 49.9 ms | 6.61 ms | 1,733 ms | 0 | run_UTpq8TWxMeGRQjxX | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=1 | 181.5 | 41.6 ms | 5.38 ms | 1,413.1 ms | 0 | run_PYAZiCJsanxKFkgf |
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=64
- Req/s
- Output tok/s
- 6,157.7
- p95 TTFT / TPOT
- 427.9 ms / 10.55 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=16
- Req/s
- Output tok/s
- 2,082.9
- p95 TTFT / TPOT
- 121.9 ms / 7.53 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=4
- Req/s
- Output tok/s
- 725.3
- p95 TTFT / TPOT
- 42 ms / 5.39 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=2
- Req/s
- Output tok/s
- 362.6
- p95 TTFT / TPOT
- 41.5 ms / 5.38 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=128
- Req/s
- Output tok/s
- 7,982.7
- p95 TTFT / TPOT
- 3,587.3 ms / 13.63 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=32
- Req/s
- Output tok/s
- 3,570.5
- p95 TTFT / TPOT
- 219.7 ms / 8.58 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=8
- Req/s
- Output tok/s
- 1,209.9
- p95 TTFT / TPOT
- 49.9 ms / 6.61 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=1
- Req/s
- Output tok/s
- 181.5
- p95 TTFT / TPOT
- 41.6 ms / 5.38 ms