gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Workload
8192 → 256
Runs
9 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)02351.701run_yK-07d4Y5vGe87ys: 0.7 req/s, p95 TTFT 2,351.7 msrun_XSLaVc1L7v1wr5aZ: 0.59 req/s, p95 TTFT 1,769 msrun_kG55hY85vGbhzWhN: 0.46 req/s, p95 TTFT 1,186.6 msrun_Y24WiWY3H_uECWVA: 0.29 req/s, p95 TTFT 610.2 msrun_z8t4rhaSqIWGcZof: 0.29 req/s, p95 TTFT 610.8 msrun_YD-0nGHOrlglf8qf: 0.29 req/s, p95 TTFT 610.5 msrun_PT98anCagtktLf9Y: 0.29 req/s, p95 TTFT 610.2 msrun_zfUui30GNOyRV1S6: 0.29 req/s, p95 TTFT 610.1 msrun_ACIJykREf1YrdTgz: 0.29 req/s, p95 TTFT 610.4 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=4180.22,351.7 ms19.78 ms0run_yK-07d4Y5vGe87ys
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=3151.61,769 ms17.25 ms0run_XSLaVc1L7v1wr5aZ
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=2117.11,186.6 ms14.61 ms0run_kG55hY85vGbhzWhN
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=173.1610.2 ms11.38 ms0run_Y24WiWY3H_uECWVA
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=173.2610.8 ms11.38 ms0run_z8t4rhaSqIWGcZof
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=173610.5 ms11.4 ms0run_YD-0nGHOrlglf8qf
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=173.3610.2 ms11.37 ms0run_PT98anCagtktLf9Y
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=173.3610.1 ms11.37 ms0run_zfUui30GNOyRV1S6
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=173.1610.4 ms11.4 ms0run_ACIJykREf1YrdTgz
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
180.2
p95 TTFT / TPOT
2,351.7 ms / 19.78 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=3
Req/s
Output tok/s
151.6
p95 TTFT / TPOT
1,769 ms / 17.25 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
117.1
p95 TTFT / TPOT
1,186.6 ms / 14.61 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
73.1
p95 TTFT / TPOT
610.2 ms / 11.38 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
73.2
p95 TTFT / TPOT
610.8 ms / 11.38 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
73
p95 TTFT / TPOT
610.5 ms / 11.4 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
73.3
p95 TTFT / TPOT
610.2 ms / 11.37 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
73.3
p95 TTFT / TPOT
610.1 ms / 11.37 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
73.1
p95 TTFT / TPOT
610.4 ms / 11.4 ms