nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Hardware
4× NVIDIA GeForce RTX 5090
Workload
1024 → 256
Runs
8 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)05167.0029run_mfqbwBkSDDFoKOa3: 25.44 req/s, p95 TTFT 617.3 msrun_qWxZVLWWpoRHNNAV: 12.62 req/s, p95 TTFT 204 msrun_3rdXHqlbZggWv4hH: 5.85 req/s, p95 TTFT 53.5 msrun_hDHgO-XcKjZB9yaq: 2.93 req/s, p95 TTFT 53.5 msrun_DapqPLH83TG71M1Y: 28.99 req/s, p95 TTFT 5,167 msrun_KyMqf9qjbzdX1J7S: 17.13 req/s, p95 TTFT 329.7 msrun_bOQg5EwiyS1jpvmn: 8.4 req/s, p95 TTFT 109.8 msrun_mHbNqg449CsCxEIt: 1.47 req/s, p95 TTFT 53 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=646,513.6617.3 ms10.63 ms2,931.1 ms0run_mfqbwBkSDDFoKOa3
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=163,231.6204 ms5.12 ms1,420.6 ms0run_qWxZVLWWpoRHNNAV
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=41,498.553.5 ms2.48 ms685.5 ms0run_3rdXHqlbZggWv4hH
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=2750.153.5 ms2.48 ms684.4 ms0run_hDHgO-XcKjZB9yaq
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=1287,420.95,167 ms9.73 ms6,802.6 ms0run_DapqPLH83TG71M1Y
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=324,385.9329.7 ms7.69 ms2,234.2 ms0run_KyMqf9qjbzdX1J7S
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=82,150.7109.8 ms3.7 ms1,120.8 ms0run_bOQg5EwiyS1jpvmn
vllm 0.27.1dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4c=1375.453 ms2.48 ms684.4 ms0run_mHbNqg449CsCxEIt
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=64
Req/s
Output tok/s
6,513.6
p95 TTFT / TPOT
617.3 ms / 10.63 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=16
Req/s
Output tok/s
3,231.6
p95 TTFT / TPOT
204 ms / 5.12 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=4
Req/s
Output tok/s
1,498.5
p95 TTFT / TPOT
53.5 ms / 2.48 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=2
Req/s
Output tok/s
750.1
p95 TTFT / TPOT
53.5 ms / 2.48 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=128
Req/s
Output tok/s
7,420.9
p95 TTFT / TPOT
5,167 ms / 9.73 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=32
Req/s
Output tok/s
4,385.9
p95 TTFT / TPOT
329.7 ms / 7.69 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=8
Req/s
Output tok/s
2,150.7
p95 TTFT / TPOT
109.8 ms / 3.7 ms
Runtime
vllm 0.27.1
Material parameters
dtype=bfloat16 max_num_seqs=87 max_model_len=16384 kv_cache_dtype=fp8 prefix_caching=false gpu_memory_utilization=0.90 max_num_batched_tokens=16384 replicas=4
Test load
c=1
Req/s
Output tok/s
375.4
p95 TTFT / TPOT
53 ms / 2.48 ms