gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Workload
256 → 1024
Runs
10 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)0140.101run_FYzLUPDhk2Yc4ZFS: 0.64 req/s, p95 TTFT 139.8 msrun_qT-6TMkmPwMdxd6z: 0.64 req/s, p95 TTFT 140.1 msrun_BzUJKLOpT1foCQmQ: 0.64 req/s, p95 TTFT 139.9 msrun_OcdPjiCmmOszJW6_: 0.48 req/s, p95 TTFT 109.9 msrun_PyERufoSv0dDFq0g: 0.31 req/s, p95 TTFT 76 msrun_PcbLolgZAS2mVFJt: 0.16 req/s, p95 TTFT 48.5 msrun_jjgZ3n3VKgVVSkgA: 0.09 req/s, p95 TTFT 38.1 msrun_riEWQzTB5E6iHuP9: 0.09 req/s, p95 TTFT 35.8 msrun_s05fh5tQpBVA9sI7: 0.09 req/s, p95 TTFT 37.7 msrun_rmh6e3v_PZG5BjWM: 0.09 req/s, p95 TTFT 37.8 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=8656.6139.8 ms12.07 ms0run_FYzLUPDhk2Yc4ZFS
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=8656.3140.1 ms12.08 ms0run_qT-6TMkmPwMdxd6z
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=8656.4139.9 ms12.08 ms0run_BzUJKLOpT1foCQmQ
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=6487109.9 ms11.85 ms0run_OcdPjiCmmOszJW6_
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=4319.576 ms12.47 ms0run_PyERufoSv0dDFq0g
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=2164.148.5 ms12.17 ms0run_PcbLolgZAS2mVFJt
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=188.938.1 ms11.22 ms0run_jjgZ3n3VKgVVSkgA
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=189.435.8 ms11.19 ms0run_riEWQzTB5E6iHuP9
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=189.437.7 ms11.19 ms0run_s05fh5tQpBVA9sI7
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=189.237.8 ms11.22 ms0run_rmh6e3v_PZG5BjWM
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
656.6
p95 TTFT / TPOT
139.8 ms / 12.07 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
656.3
p95 TTFT / TPOT
140.1 ms / 12.08 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
656.4
p95 TTFT / TPOT
139.9 ms / 12.08 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=6
Req/s
Output tok/s
487
p95 TTFT / TPOT
109.9 ms / 11.85 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
319.5
p95 TTFT / TPOT
76 ms / 12.47 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
164.1
p95 TTFT / TPOT
48.5 ms / 12.17 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
88.9
p95 TTFT / TPOT
38.1 ms / 11.22 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
89.4
p95 TTFT / TPOT
35.8 ms / 11.19 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
89.4
p95 TTFT / TPOT
37.7 ms / 11.19 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
89.2
p95 TTFT / TPOT
37.8 ms / 11.22 ms