- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 8192 → 256
- Runs
- 6 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=8 | 282.5 | 3,139.1 ms | 26.21 ms | 7,775.3 ms | 0 | run_Dr7IbnE-Db45CNX1 | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=4 | 203.2 | 2,102.7 ms | 17.59 ms | 5,552.7 ms | 0 | run_rVGmECPgjdGY0zL0 | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=32 | 390.6 | 14,271.1 ms | 75.75 ms | 30,508.9 ms | 0 | run_hAASlU_O0xen7oUo | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=2 | 129.4 | 1,060.2 ms | 13.39 ms | 3,961.3 ms | 0 | run_cUEhn7RgcgZxtzFe | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=1 | 74.7 | 541.4 ms | 11.38 ms | 3,439.2 ms | 0 | run__BycYOCjzrAsfrwa | |
| vllm 0.23.0+cu129 | dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor | c=16 | 350.4 | 5,835.3 ms | 43.49 ms | 13,222.6 ms | 0 | run_HubfBjV6uSNZsQwT |
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=8
- Req/s
- Output tok/s
- 282.5
- p95 TTFT / TPOT
- 3,139.1 ms / 26.21 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=4
- Req/s
- Output tok/s
- 203.2
- p95 TTFT / TPOT
- 2,102.7 ms / 17.59 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=32
- Req/s
- Output tok/s
- 390.6
- p95 TTFT / TPOT
- 14,271.1 ms / 75.75 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=2
- Req/s
- Output tok/s
- 129.4
- p95 TTFT / TPOT
- 1,060.2 ms / 13.39 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=1
- Req/s
- Output tok/s
- 74.7
- p95 TTFT / TPOT
- 541.4 ms / 11.38 ms
- Runtime
- vllm 0.23.0+cu129
- Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor- Test load
- c=16
- Req/s
- Output tok/s
- 350.4
- p95 TTFT / TPOT
- 5,835.3 ms / 43.49 ms