nvidia/Gemma-4-26B-A4B-NVFP4

Hardware
4× NVIDIA GeForce RTX 5090
Workload
1024 → 256
Runs
8 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)03587.3031run_vR-aN6dQN7oN9vEX: 24.05 req/s, p95 TTFT 427.9 msrun_aFn0Jx0E_Ml6eRgq: 8.14 req/s, p95 TTFT 121.9 msrun_TLgC2HLVcevQOitz: 2.83 req/s, p95 TTFT 42 msrun_2Hz338t5HaGnvCAY: 1.42 req/s, p95 TTFT 41.5 msrun_H2PYudiZpc26_sUQ: 31.18 req/s, p95 TTFT 3,587.3 msrun_ELXdCg4-RIo_ojeG: 13.95 req/s, p95 TTFT 219.7 msrun_UTpq8TWxMeGRQjxX: 4.73 req/s, p95 TTFT 49.9 msrun_PYAZiCJsanxKFkgf: 0.71 req/s, p95 TTFT 41.6 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=646,157.7427.9 ms10.55 ms2,895.7 ms0run_vR-aN6dQN7oN9vEX
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=162,082.9121.9 ms7.53 ms1,977.9 ms0run_aFn0Jx0E_Ml6eRgq
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=4725.342 ms5.39 ms1,415.4 ms0run_TLgC2HLVcevQOitz
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=2362.641.5 ms5.38 ms1,413.2 ms0run_2Hz338t5HaGnvCAY
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=1287,982.73,587.3 ms13.63 ms6,850.7 ms0run_H2PYudiZpc26_sUQ
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=323,570.5219.7 ms8.58 ms2,282 ms0run_ELXdCg4-RIo_ojeG
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=81,209.949.9 ms6.61 ms1,733 ms0run_UTpq8TWxMeGRQjxX
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=1181.541.6 ms5.38 ms1,413.1 ms0run_PYAZiCJsanxKFkgf
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=64
Req/s
Output tok/s
6,157.7
p95 TTFT / TPOT
427.9 ms / 10.55 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=16
Req/s
Output tok/s
2,082.9
p95 TTFT / TPOT
121.9 ms / 7.53 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=4
Req/s
Output tok/s
725.3
p95 TTFT / TPOT
42 ms / 5.39 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=2
Req/s
Output tok/s
362.6
p95 TTFT / TPOT
41.5 ms / 5.38 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=128
Req/s
Output tok/s
7,982.7
p95 TTFT / TPOT
3,587.3 ms / 13.63 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=32
Req/s
Output tok/s
3,570.5
p95 TTFT / TPOT
219.7 ms / 8.58 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=8
Req/s
Output tok/s
1,209.9
p95 TTFT / TPOT
49.9 ms / 6.61 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=1
Req/s
Output tok/s
181.5
p95 TTFT / TPOT
41.6 ms / 5.38 ms