google/gemma-4-12B-it

Hardware
1× NVIDIA GeForce RTX 5090
Workload
8192 → 256
Runs
6 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)014271.102run_Dr7IbnE-Db45CNX1: 1.1 req/s, p95 TTFT 3,139.1 msrun_rVGmECPgjdGY0zL0: 0.79 req/s, p95 TTFT 2,102.7 msrun_hAASlU_O0xen7oUo: 1.53 req/s, p95 TTFT 14,271.1 msrun_cUEhn7RgcgZxtzFe: 0.51 req/s, p95 TTFT 1,060.2 msrun__BycYOCjzrAsfrwa: 0.29 req/s, p95 TTFT 541.4 msrun_HubfBjV6uSNZsQwT: 1.37 req/s, p95 TTFT 5,835.3 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=8282.53,139.1 ms26.21 ms7,775.3 ms0run_Dr7IbnE-Db45CNX1
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=4203.22,102.7 ms17.59 ms5,552.7 ms0run_rVGmECPgjdGY0zL0
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=32390.614,271.1 ms75.75 ms30,508.9 ms0run_hAASlU_O0xen7oUo
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=2129.41,060.2 ms13.39 ms3,961.3 ms0run_cUEhn7RgcgZxtzFe
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=174.7541.4 ms11.38 ms3,439.2 ms0run__BycYOCjzrAsfrwa
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=16350.45,835.3 ms43.49 ms13,222.6 ms0run_HubfBjV6uSNZsQwT
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=8
Req/s
Output tok/s
282.5
p95 TTFT / TPOT
3,139.1 ms / 26.21 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=4
Req/s
Output tok/s
203.2
p95 TTFT / TPOT
2,102.7 ms / 17.59 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=32
Req/s
Output tok/s
390.6
p95 TTFT / TPOT
14,271.1 ms / 75.75 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=2
Req/s
Output tok/s
129.4
p95 TTFT / TPOT
1,060.2 ms / 13.39 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=1
Req/s
Output tok/s
74.7
p95 TTFT / TPOT
541.4 ms / 11.38 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=16
Req/s
Output tok/s
350.4
p95 TTFT / TPOT
5,835.3 ms / 43.49 ms