gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal
Workload
256 → 1024
Runs
10 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)01235.802run_ucnXL4xlzYB0WExx: 2.27 req/s, p95 TTFT 1,235.8 msrun_JEq9OLfaQCO9CHzn: 1.94 req/s, p95 TTFT 117.7 msrun_l27XVarmQiKyGJf9: 1.94 req/s, p95 TTFT 117.4 msrun_b9VV0WgkqDwEwZ1z: 1.87 req/s, p95 TTFT 109.8 msrun_ZZ4xMERgyAoHQgGr: 1.49 req/s, p95 TTFT 91.1 msrun_Y8t06MkjhQ5Hbobt: 0.83 req/s, p95 TTFT 58.1 msrun_pjgVwVQ4fyoDIT9Y: 0.43 req/s, p95 TTFT 43.6 msrun_Y055cveCFjmgKavh: 0.43 req/s, p95 TTFT 43.6 msrun_pVFDaY3ZKHpMkDRJ: 0.43 req/s, p95 TTFT 43.6 msrun_5zWNumbaU3hXyqUw: 0.43 req/s, p95 TTFT 43.7 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=82,329.11,235.8 ms5.24 ms0run_ucnXL4xlzYB0WExx
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=61,986.2117.7 ms6.08 ms0run_JEq9OLfaQCO9CHzn
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=61,986.7117.4 ms6.08 ms0run_l27XVarmQiKyGJf9
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=61,914.7109.8 ms4.44 ms0run_b9VV0WgkqDwEwZ1z
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=41,52791.1 ms3.18 ms0run_ZZ4xMERgyAoHQgGr
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=2854.358.1 ms4.06 ms0run_Y8t06MkjhQ5Hbobt
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=1436.743.6 ms2.84 ms0run_pjgVwVQ4fyoDIT9Y
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1436.743.6 ms2.84 ms0run_Y055cveCFjmgKavh
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1436.743.6 ms2.84 ms0run_pVFDaY3ZKHpMkDRJ
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1435.843.7 ms2.85 ms0run_5zWNumbaU3hXyqUw
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
2,329.1
p95 TTFT / TPOT
1,235.8 ms / 5.24 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=6
Req/s
Output tok/s
1,986.2
p95 TTFT / TPOT
117.7 ms / 6.08 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=6
Req/s
Output tok/s
1,986.7
p95 TTFT / TPOT
117.4 ms / 6.08 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=6
Req/s
Output tok/s
1,914.7
p95 TTFT / TPOT
109.8 ms / 4.44 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
1,527
p95 TTFT / TPOT
91.1 ms / 3.18 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
854.3
p95 TTFT / TPOT
58.1 ms / 4.06 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
436.7
p95 TTFT / TPOT
43.6 ms / 2.84 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
436.7
p95 TTFT / TPOT
43.6 ms / 2.84 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
436.7
p95 TTFT / TPOT
43.6 ms / 2.84 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
435.8
p95 TTFT / TPOT
43.7 ms / 2.85 ms