- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 256 → 1024
- Runs
- 6 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=8 | 698.3 | 139.9 ms | 11.4 ms | 11,769.7 ms | 0 | run_vfoxggT7OGVrNnNz | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=4 | 361.5 | 96.2 ms | 11.06 ms | 11,406.3 ms | 0 | run_nHtJhvJi4NcGvPwh | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=32 | 2,269.6 | 474 ms | 14.02 ms | 14,446.5 ms | 0 | run_Nw1U2fDeSwJfLpPp | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=2 | 185.4 | 63 ms | 10.77 ms | 11,060.5 ms | 0 | run_JRoYK-1Zc-F0gKvw | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=1 | 93.4 | 27.4 ms | 10.71 ms | 10,984 ms | 0 | run_JoZqHLjbBj8U5Nft | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=16 | 1,297.6 | 257.2 ms | 12.26 ms | 12,641.2 ms | 0 | run_tFo55ZNrx3T150mL |
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=8
- Req/s
- Output tok/s
- 698.3
- p95 TTFT / TPOT
- 139.9 ms / 11.4 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=4
- Req/s
- Output tok/s
- 361.5
- p95 TTFT / TPOT
- 96.2 ms / 11.06 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=32
- Req/s
- Output tok/s
- 2,269.6
- p95 TTFT / TPOT
- 474 ms / 14.02 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=2
- Req/s
- Output tok/s
- 185.4
- p95 TTFT / TPOT
- 63 ms / 10.77 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=1
- Req/s
- Output tok/s
- 93.4
- p95 TTFT / TPOT
- 27.4 ms / 10.71 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=16
- Req/s
- Output tok/s
- 1,297.6
- p95 TTFT / TPOT
- 257.2 ms / 12.26 ms