gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Workload
4096 → 1
Runs
7 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)02209.904run_FxKGqjdkjx14BwR-: 3.74 req/s, p95 TTFT 2,209.9 msrun_JMb-m604g1F6rFj6: 3.73 req/s, p95 TTFT 1,668.4 msrun_aMD1tPYsFLz_1wVn: 3.74 req/s, p95 TTFT 1,131.7 msrun_PxlhuqS2TEpaqDA1: 3.74 req/s, p95 TTFT 595.7 msrun_OJCpIlKXcUAitTF9: 3.74 req/s, p95 TTFT 595 msrun_xitx0BeMs3KHD2V7: 3.73 req/s, p95 TTFT 596.2 msrun_8CkbuHSlqfBM8f0X: 3.65 req/s, p95 TTFT 275.7 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=83.72,209.9 ms0 ms0run_FxKGqjdkjx14BwR-
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=63.71,668.4 ms0 ms0run_JMb-m604g1F6rFj6
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=43.71,131.7 ms0 ms0run_aMD1tPYsFLz_1wVn
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=23.7595.7 ms0 ms0run_PxlhuqS2TEpaqDA1
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=23.7595 ms0 ms0run_OJCpIlKXcUAitTF9
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=23.7596.2 ms0 ms0run_xitx0BeMs3KHD2V7
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=13.6275.7 ms0 ms0run_8CkbuHSlqfBM8f0X
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
3.7
p95 TTFT / TPOT
2,209.9 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=6
Req/s
Output tok/s
3.7
p95 TTFT / TPOT
1,668.4 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
3.7
p95 TTFT / TPOT
1,131.7 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
3.7
p95 TTFT / TPOT
595.7 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
3.7
p95 TTFT / TPOT
595 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
3.7
p95 TTFT / TPOT
596.2 ms / 0 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
3.6
p95 TTFT / TPOT
275.7 ms / 0 ms