google/gemma-4-12B-it

Hardware
1× NVIDIA GeForce RTX 5090
Workload
4096 → 1
Runs
7 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)015701.004run_MLQL2JZYOrxrOWI4: 4.06 req/s, p95 TTFT 1,971 msrun_wcDKJV0Dw36T303u: 4.07 req/s, p95 TTFT 15,701 msrun_kI4buRXw31qtTUFw: 4.06 req/s, p95 TTFT 985.4 msrun_cgv1tBBK1i6GGRZ_: 4.07 req/s, p95 TTFT 7,847.7 msrun_h6nca3C4GMKr_MPQ: 4.02 req/s, p95 TTFT 498.7 msrun_8pl8xnDbAwa8IvHO: 3.9 req/s, p95 TTFT 259.4 msrun_g7f5L-ATmM_vBrcg: 4.07 req/s, p95 TTFT 3,927.7 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=84.11,971 ms0 ms1,971 ms0run_MLQL2JZYOrxrOWI4
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=644.115,701 ms0 ms15,701 ms0run_wcDKJV0Dw36T303u
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=44.1985.4 ms0 ms985.4 ms0run_kI4buRXw31qtTUFw
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=324.17,847.7 ms0 ms7,847.7 ms0run_cgv1tBBK1i6GGRZ_
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=24498.7 ms0 ms498.7 ms0run_h6nca3C4GMKr_MPQ
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=13.9259.4 ms0 ms259.4 ms0run_8pl8xnDbAwa8IvHO
vllm 0.23.0+cu129dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensorc=164.13,927.7 ms0 ms3,927.7 ms0run_g7f5L-ATmM_vBrcg
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=8
Req/s
Output tok/s
4.1
p95 TTFT / TPOT
1,971 ms / 0 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=64
Req/s
Output tok/s
4.1
p95 TTFT / TPOT
15,701 ms / 0 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=4
Req/s
Output tok/s
4.1
p95 TTFT / TPOT
985.4 ms / 0 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=32
Req/s
Output tok/s
4.1
p95 TTFT / TPOT
7,847.7 ms / 0 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=2
Req/s
Output tok/s
4
p95 TTFT / TPOT
498.7 ms / 0 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=1
Req/s
Output tok/s
3.9
p95 TTFT / TPOT
259.4 ms / 0 ms
Runtime
vllm 0.23.0+cu129
Material parameters
dtype=bfloat16 max_num_seqs=64 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.92 max_num_batched_tokens=16384 runtime_weight_quantization=fp8_per_tensor
Test load
c=16
Req/s
Output tok/s
4.1
p95 TTFT / TPOT
3,927.7 ms / 0 ms