nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash
Workload
1024 → 256
Runs
6 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)0529.309run_KPlG6ospa82y9K0I: 7.18 req/s, p95 TTFT 176.1 msrun_G_VhMGaPBYxOstc_: 4.77 req/s, p95 TTFT 96.3 msrun_KUwRa7hzWEZqwMkg: 1.97 req/s, p95 TTFT 48.2 msrun_dhCcdGR7Ab_02vHG: 9.49 req/s, p95 TTFT 523.7 msrun_sooclCxdf5Rwfmlt: 9.3 req/s, p95 TTFT 522.7 msrun_rtUGFh2HoRtyVSzF: 9.25 req/s, p95 TTFT 529.3 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=81,838.5176.1 ms11.59 ms3,043.9 ms0run_KPlG6ospa82y9K0I
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=41,221.596.3 ms8.76 ms2,294.6 ms0run_G_VhMGaPBYxOstc_
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=1504.648.2 ms5.53 ms1,455.1 ms0run_KUwRa7hzWEZqwMkg
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=162,429.3523.7 ms16.87 ms4,416.4 ms0run_dhCcdGR7Ab_02vHG
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=162,381.9522.7 ms16.46 ms4,338.2 ms0run_sooclCxdf5Rwfmlt
vllm 0.28.0DFlash · NVFP4dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384c=162,366.9529.3 ms17.09 ms4,439.8 ms0run_rtUGFh2HoRtyVSzF
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=8
Req/s
Output tok/s
1,838.5
p95 TTFT / TPOT
176.1 ms / 11.59 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=4
Req/s
Output tok/s
1,221.5
p95 TTFT / TPOT
96.3 ms / 8.76 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=1
Req/s
Output tok/s
504.6
p95 TTFT / TPOT
48.2 ms / 5.53 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
2,429.3
p95 TTFT / TPOT
523.7 ms / 16.87 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
2,381.9
p95 TTFT / TPOT
522.7 ms / 16.46 ms
Runtime
vllm 0.28.0DFlash · NVFP4 draft
Material parameters
dtype=bfloat16 max_num_seqs=32 enforce_eager=false max_model_len=9216 kv_cache_dtype=fp8 prefix_caching=false chunked_prefill=true expert_parallel=false async_scheduling=true gpu_memory_utilization=0.88 max_num_batched_tokens=16384
Test load
c=16
Req/s
Output tok/s
2,366.9
p95 TTFT / TPOT
529.3 ms / 17.09 ms