- Hardware
- 4× NVIDIA GeForce RTX 5090
- Workload
- 1024 → 256
- Runs
- 8 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=64 | 2,434.6 | 5,698.2 ms | 20.12 ms | 10,957.2 ms | 0 | run_EpUX0RC2nD27Favg | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=16 | 873.6 | 356.8 ms | 17.19 ms | 4,611.9 ms | 0 | run_ZijkPzwF5OpCpk96 | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=4 | 269.5 | 95.7 ms | 14.54 ms | 3,803.2 ms | 0 | run_fnPS9SkKESxDA8RX | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=2 | 134.8 | 95 ms | 14.53 ms | 3,799.4 ms | 0 | run_arKAxEVDmDFnwMzM | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=128 | 2,421.7 | 12,221 ms | 19.32 ms | 16,634.3 ms | 0 | run_xlrn81sRTyav7HWD | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=32 | 1,559.1 | 719.3 ms | 17.59 ms | 4,914.9 ms | 0 | run_F5yp10AfL5aiaDng | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=8 | 480.4 | 178 ms | 16.23 ms | 4,270.2 ms | 0 | run_M-dDOg-ZxrerXm3Y | |
| vllm 0.27.1 | dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4 | c=1 | 67.4 | 95 ms | 14.53 ms | 3,799.3 ms | 0 | run_xm_EmAr3djdWaetd |
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=64
- Req/s
- Output tok/s
- 2,434.6
- p95 TTFT / TPOT
- 5,698.2 ms / 20.12 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=16
- Req/s
- Output tok/s
- 873.6
- p95 TTFT / TPOT
- 356.8 ms / 17.19 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=4
- Req/s
- Output tok/s
- 269.5
- p95 TTFT / TPOT
- 95.7 ms / 14.54 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=2
- Req/s
- Output tok/s
- 134.8
- p95 TTFT / TPOT
- 95 ms / 14.53 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=128
- Req/s
- Output tok/s
- 2,421.7
- p95 TTFT / TPOT
- 12,221 ms / 19.32 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=32
- Req/s
- Output tok/s
- 1,559.1
- p95 TTFT / TPOT
- 719.3 ms / 17.59 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=8
- Req/s
- Output tok/s
- 480.4
- p95 TTFT / TPOT
- 178 ms / 16.23 ms
- Runtime
- vllm 0.27.1
- Material parameters
dtype=bfloat16 max_num_seqs=26 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4- Test load
- c=1
- Req/s
- Output tok/s
- 67.4
- p95 TTFT / TPOT
- 95 ms / 14.53 ms