mistralai/Mistral-Small-4-119B-2603-NVFP4

Hardware
4× NVIDIA GeForce RTX 5090
Workload
1024 → 256
Runs
8 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)04129.3012run_lPHnsM7weInxjd_C: 9.09 req/s, p95 TTFT 2,528.4 msrun_QjI5OdTC3LlWDCBs: 5.19 req/s, p95 TTFT 652.7 msrun_8FxfHM5XZUC0ZUa6: 2.41 req/s, p95 TTFT 194 msrun_QSDg-z_HDZUQlhfi: 1.26 req/s, p95 TTFT 109.3 msrun_gfWaR-EN7-iG8d8f: 11.79 req/s, p95 TTFT 4,129.3 msrun_F_ID-98IPqyMT3Zc: 4.99 req/s, p95 TTFT 1,269 msrun_ExfD7Ne8gmqWHZBf: 3.69 req/s, p95 TTFT 343.9 msrun_0fdYMWzrl34_j431: 0.92 req/s, p95 TTFT 58.1 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=642,327.42,528.4 ms25.04 ms7,674.9 ms0run_lPHnsM7weInxjd_C
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=161,329.5652.7 ms11.48 ms3,069.3 ms0run_QjI5OdTC3LlWDCBs
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=4617.7194 ms6.26 ms1,662.1 ms0run_8FxfHM5XZUC0ZUa6
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=2323.3109.3 ms5.98 ms1,585.2 ms0run_QSDg-z_HDZUQlhfi
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=1283,018.54,129.3 ms40.97 ms13,044.9 ms0run_gfWaR-EN7-iG8d8f
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=321,277.31,269 ms59.17 ms15,730.2 ms0run_F_ID-98IPqyMT3Zc
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=8944.6343.9 ms8.25 ms2,179.6 ms0run_ExfD7Ne8gmqWHZBf
vllm 0.27.1dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4c=123558.1 ms4.05 ms1,090 ms0run_0fdYMWzrl34_j431
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=64
Req/s
Output tok/s
2,327.4
p95 TTFT / TPOT
2,528.4 ms / 25.04 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=16
Req/s
Output tok/s
1,329.5
p95 TTFT / TPOT
652.7 ms / 11.48 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=4
Req/s
Output tok/s
617.7
p95 TTFT / TPOT
194 ms / 6.26 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=2
Req/s
Output tok/s
323.3
p95 TTFT / TPOT
109.3 ms / 5.98 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=128
Req/s
Output tok/s
3,018.5
p95 TTFT / TPOT
4,129.3 ms / 40.97 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=32
Req/s
Output tok/s
1,277.3
p95 TTFT / TPOT
1,269 ms / 59.17 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=8
Req/s
Output tok/s
944.6
p95 TTFT / TPOT
343.9 ms / 8.25 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=128 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 tensor=4
Test load
c=1
Req/s
Output tok/s
235
p95 TTFT / TPOT
58.1 ms / 4.05 ms