gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
incoai/Qwen3.8-27B-DFlash2
Workload
256 → 1024
Runs
10 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)04179.501run_ZdaS_isTyCbrlzky: 1.27 req/s, p95 TTFT 4,179.5 msrun_YSoYEmpUtwFOGqIT: 1.27 req/s, p95 TTFT 2,425.1 msrun_hqnqmcX5ovC9wOgL: 1.28 req/s, p95 TTFT 86.6 msrun_zsR6FSLwBSv6O6Xy: 1.28 req/s, p95 TTFT 86.8 msrun_2TUVAgymXwzhJUp2: 1.27 req/s, p95 TTFT 89.3 msrun_S_8hksOANJ-0LjMh: 0.74 req/s, p95 TTFT 60.5 msrun_3gDbZGc0Xjt8ZzwE: 0.41 req/s, p95 TTFT 44.4 msrun_dkYZh2AwUTjESR_A: 0.41 req/s, p95 TTFT 44.5 msrun_xCjQi0gT6Va71S5g: 0.41 req/s, p95 TTFT 44.3 msrun_fHL70cqgxKpV-z06: 0.41 req/s, p95 TTFT 44.4 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=81,305.14,179.5 ms5.69 ms0run_ZdaS_isTyCbrlzky
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=61,3052,425.1 ms5.69 ms0run_YSoYEmpUtwFOGqIT
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=41,305.786.6 ms5.68 ms0run_hqnqmcX5ovC9wOgL
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=41,30686.8 ms5.68 ms0run_zsR6FSLwBSv6O6Xy
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=41,30589.3 ms5.69 ms0run_2TUVAgymXwzhJUp2
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=2761.560.5 ms3.85 ms0run_S_8hksOANJ-0LjMh
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=1414.944.4 ms3.02 ms0run_3gDbZGc0Xjt8ZzwE
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1418.844.5 ms2.99 ms0run_dkYZh2AwUTjESR_A
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1418.944.3 ms2.99 ms0run_xCjQi0gT6Va71S5g
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1418.944.4 ms2.99 ms0run_fHL70cqgxKpV-z06
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=8
Req/s
Output tok/s
1,305.1
p95 TTFT / TPOT
4,179.5 ms / 5.69 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=6
Req/s
Output tok/s
1,305
p95 TTFT / TPOT
2,425.1 ms / 5.69 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=4
Req/s
Output tok/s
1,305.7
p95 TTFT / TPOT
86.6 ms / 5.68 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=4
Req/s
Output tok/s
1,306
p95 TTFT / TPOT
86.8 ms / 5.68 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=4
Req/s
Output tok/s
1,305
p95 TTFT / TPOT
89.3 ms / 5.69 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=2
Req/s
Output tok/s
761.5
p95 TTFT / TPOT
60.5 ms / 3.85 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=1
Req/s
Output tok/s
414.9
p95 TTFT / TPOT
44.4 ms / 3.02 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
418.8
p95 TTFT / TPOT
44.5 ms / 2.99 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
418.9
p95 TTFT / TPOT
44.3 ms / 2.99 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
418.9
p95 TTFT / TPOT
44.4 ms / 2.99 ms