RedHatAI/gemma-4-31B-it-NVFP4

Hardware
1× NVIDIA GeForce RTX 5090
Decoding method
Speculative decoding
Workload
256 → 1024
Runs
6 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)05010.701run_OpmYrrUSfdj2ebxm: 0.56 req/s, p95 TTFT 101.1 msrun_3fk4K9Ig4bfl7k32: 0.56 req/s, p95 TTFT 130.8 msrun_UeB5ytUp-kIr2xJ1: 0.56 req/s, p95 TTFT 98.1 msrun_oibdkF5ODm6f1LiS: 0.29 req/s, p95 TTFT 64.6 msrun_Ax841n8vtQ-5QCL8: 0.08 req/s, p95 TTFT 46.6 msrun_CUB80RbpKM0BL_6P: 0.82 req/s, p95 TTFT 5,010.7 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.28.0Speculativedtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=8576.6101.1 ms17.84 ms18,314.2 ms0run_OpmYrrUSfdj2ebxm
vllm 0.28.0Speculativedtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=8575.2130.8 ms17.83 ms18,306.6 ms0run_3fk4K9Ig4bfl7k32
vllm 0.28.0Speculativedtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=8578.398.1 ms17.79 ms18,266 ms0run_UeB5ytUp-kIr2xJ1
vllm 0.28.0Speculativedtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=4294.864.6 ms17.47 ms17,933.7 ms0run_oibdkF5ODm6f1LiS
vllm 0.28.0Speculativedtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=180.846.6 ms16.56 ms16,970.6 ms0run_Ax841n8vtQ-5QCL8
vllm 0.28.0Speculativedtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=168375,010.7 ms25.1 ms26,407.8 ms0run_CUB80RbpKM0BL_6P
Runtime
vllm 0.28.0Speculative decoding
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
576.6
p95 TTFT / TPOT
101.1 ms / 17.84 ms
Runtime
vllm 0.28.0Speculative decoding
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
575.2
p95 TTFT / TPOT
130.8 ms / 17.83 ms
Runtime
vllm 0.28.0Speculative decoding
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
578.3
p95 TTFT / TPOT
98.1 ms / 17.79 ms
Runtime
vllm 0.28.0Speculative decoding
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=4
Req/s
Output tok/s
294.8
p95 TTFT / TPOT
64.6 ms / 17.47 ms
Runtime
vllm 0.28.0Speculative decoding
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
80.8
p95 TTFT / TPOT
46.6 ms / 16.56 ms
Runtime
vllm 0.28.0Speculative decoding
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
837
p95 TTFT / TPOT
5,010.7 ms / 25.1 ms