gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal
Workload
4096 → 1
Runs
7 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)02345.404run_xeMd7hTQ5EwuloOC: 3.51 req/s, p95 TTFT 2,345.4 msrun_xfPAVyBvYPhyO-dq: 3.51 req/s, p95 TTFT 1,715.1 msrun_l_6LlrJB8HmNgEjk: 3.52 req/s, p95 TTFT 1,198.5 msrun_0zHkGCRVQPkbxlMo: 3.52 req/s, p95 TTFT 629.5 msrun_p3J3PPZDQKR2lQbw: 3.52 req/s, p95 TTFT 629.4 msrun_OeJ0VW4SlzGzEQ9M: 3.51 req/s, p95 TTFT 630 msrun_ZME1GlfTDdYp3UMt: 3.42 req/s, p95 TTFT 294.6 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=83.52,345.4 ms0 ms0run_xeMd7hTQ5EwuloOC
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=63.51,715.1 ms0 ms0run_xfPAVyBvYPhyO-dq
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=43.51,198.5 ms0 ms0run_l_6LlrJB8HmNgEjk
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=23.5629.5 ms0 ms0run_0zHkGCRVQPkbxlMo
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=23.5629.4 ms0 ms0run_p3J3PPZDQKR2lQbw
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=23.5630 ms0 ms0run_OeJ0VW4SlzGzEQ9M
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=13.4294.6 ms0 ms0run_ZME1GlfTDdYp3UMt
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
3.5
p95 TTFT / TPOT
2,345.4 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=6
Req/s
Output tok/s
3.5
p95 TTFT / TPOT
1,715.1 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
3.5
p95 TTFT / TPOT
1,198.5 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
3.5
p95 TTFT / TPOT
629.5 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
3.5
p95 TTFT / TPOT
629.4 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
3.5
p95 TTFT / TPOT
630 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
3.4
p95 TTFT / TPOT
294.6 ms / 0 ms