gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal
Workload
8192 → 256
Runs
9 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)02484.101run_o9RGP_25BQeQR59o: 1.19 req/s, p95 TTFT 2,484.1 msrun_3BDbEa4p7fG_ces3: 1.19 req/s, p95 TTFT 1,302 msrun_baQvEtRupq-RH4yj: 1.07 req/s, p95 TTFT 1,186.1 msrun_j7-68fHeePgq1RaT: 0.81 req/s, p95 TTFT 648.6 msrun_RP8_yJMtqc-rJRKu: 0.81 req/s, p95 TTFT 647.5 msrun_wI6gPdjut78iyVVR: 0.81 req/s, p95 TTFT 649.9 msrun_qtIM7sYO5VUQ2HmX: 0.81 req/s, p95 TTFT 648.6 msrun_M7cqO7Ecb-elWHRy: 0.81 req/s, p95 TTFT 648.7 msrun_Af-H_BUeFX_g2m6m: 0.81 req/s, p95 TTFT 651.4 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=4305.62,484.1 ms10.89 ms0run_o9RGP_25BQeQR59o
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=33041,302 ms13.43 ms0run_3BDbEa4p7fG_ces3
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=2274.21,186.1 ms7.78 ms0run_baQvEtRupq-RH4yj
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=1207.8648.6 ms3.94 ms0run_j7-68fHeePgq1RaT
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=1207.9647.5 ms3.94 ms0run_RP8_yJMtqc-rJRKu
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=1208649.9 ms3.93 ms0run_wI6gPdjut78iyVVR
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1208.1648.6 ms3.93 ms0run_qtIM7sYO5VUQ2HmX
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1208.1648.7 ms3.93 ms0run_M7cqO7Ecb-elWHRy
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1207.9651.4 ms3.94 ms0run_Af-H_BUeFX_g2m6m
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
305.6
p95 TTFT / TPOT
2,484.1 ms / 10.89 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=3
Req/s
Output tok/s
304
p95 TTFT / TPOT
1,302 ms / 13.43 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
274.2
p95 TTFT / TPOT
1,186.1 ms / 7.78 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
207.8
p95 TTFT / TPOT
648.6 ms / 3.94 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
207.9
p95 TTFT / TPOT
647.5 ms / 3.94 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
208
p95 TTFT / TPOT
649.9 ms / 3.93 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
208.1
p95 TTFT / TPOT
648.6 ms / 3.93 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
208.1
p95 TTFT / TPOT
648.7 ms / 3.93 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
207.9
p95 TTFT / TPOT
651.4 ms / 3.94 ms