nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash
Workload
256 → 1024
Runs
6 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)0292.903run_zibeZ2Nv5x9cwBq5: 2.38 req/s, p95 TTFT 119.7 msrun_-32mTiWNVl85G6QU: 1.47 req/s, p95 TTFT 69.4 msrun_2gqaWlkvthlx-A4P: 0.54 req/s, p95 TTFT 19.3 msrun_73YNwtO9ST3Ato8a: 3.5 req/s, p95 TTFT 194.3 msrun_MsQypIw6xrxOnrhb: 3.43 req/s, p95 TTFT 221.6 msrun_YU6G5_Z0LYBxsQ5g: 3.48 req/s, p95 TTFT 292.9 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=82,432.2119.7 ms8.99 ms9,256.2 ms0run_zibeZ2Nv5x9cwBq5
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=41,504.569.4 ms7.14 ms7,379 ms0run_-32mTiWNVl85G6QU
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=1556.319.3 ms5.07 ms5,206.4 ms0run_2gqaWlkvthlx-A4P
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=163,583.3194.3 ms11.35 ms11,668.9 ms0run_73YNwtO9ST3Ato8a
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=163,513.2221.6 ms11.33 ms11,658.1 ms0run_MsQypIw6xrxOnrhb
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=163,566.9292.9 ms10.86 ms11,172.7 ms0run_YU6G5_Z0LYBxsQ5g
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
2,432.2
p95 TTFT / TPOT
119.7 ms / 8.99 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=4
Req/s
Output tok/s
1,504.5
p95 TTFT / TPOT
69.4 ms / 7.14 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
556.3
p95 TTFT / TPOT
19.3 ms / 5.07 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
3,583.3
p95 TTFT / TPOT
194.3 ms / 11.35 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
3,513.2
p95 TTFT / TPOT
221.6 ms / 11.33 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
3,566.9
p95 TTFT / TPOT
292.9 ms / 10.86 ms