gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

completed
Request throughput
1.51 req/s
Output throughput
386 tok/s
p95 TTFT
83.1 ms
p95 TPOT
2.66 ms

Manifest warnings

Review these published caveats before interpreting or comparing the metrics.

  • Review imported facts against the local model, hardware, launch command, and benchmark command before publication.

Server launch command

CUDA_VISIBLE_DEVICES=0 sglang serve \
  --model-path gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --revision 0cc27958cefbbe231782ec8511de8c4eb5233348 \
  --served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --port 9000 \
  --host 127.0.0.1 \
  --dtype bfloat16 \
  --mem-fraction-static 0.9 \
  --context-length 9216 \
  --max-running-requests 1 \
  --tp 1 \
  --kv-cache-dtype fp8_e4m3 \
  --chunked-prefill-size 1024 \
  --attention-backend flashinfer \
  --cuda-graph-max-bs 1 \
  --trust-remote-code \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-full-memory-ratio 5.61 \
  --mamba-radix-cache-strategy extra_buffer_lazy \
  --mamba-ssm-dtype bfloat16 \
  --cuda-graph-bs-prefill 64 128 256 512 1024 \
  --enable-memory-saver \
  --enable-hierarchical-cache \
  --hicache-ratio 2 \
  --hicache-write-policy write_through \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 \
  --speculative-draft-model-revision dedf8df68adfb1afeaf7b7480c0a0243108177b4 \
  --speculative-draft-model-quantization unquant \
  --speculative-num-draft-tokens 8

Run details

Workload
1024 → 256
Methodology
closed loop · streaming
Test load
1 concurrent requests
Endpoint
openai-compatible-completions
Requests
16 unmeasured warmups per invocation after one readiness request.
Dataset
sha256=7f10b755bb8a97e1f56cbe8e2e66b6af3cc8078b04f2dd3474db54ddba85214b generator=runpile-sweeps-tokenizer-verified-jsonl synthetic=true generator_revision=acfe836377740b09ca2b148328ed3e98b8989baa77300bb70a8c1df9048a9199 shared_prefix_tokens=0
Sampling
mode=greedy temperature=0
Stopping
ignore_eos=true max_output_tokens=256
Prefix and cache
cache_state=warm-host cache_enabled=false shared_prefix_tokens=0
Client
name=sglang benchmark serving backend=sglang version=0.0.0.dev1+g5f55db35e placement=same-host
Hardware
1× NVIDIA GeForce RTX 5090
Runtime
sglang 0.0.0.dev1+g5f55db35e
Decoding method
DFlash2 · BF16 draft · 8 draft tokens
Draft model
incoai/Qwen3.8-27B-DFlash2 · BF16revision dedf8df68adfb1afeaf7b7480c0a0243108177b4
Runtime configuration
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Measurement
sglang benchmark serving 0.0.0.dev1+g5f55db35e
Timing
client-observed streaming request latency · client overhead included
Repetitions
1
Failure handling
Failed requests counted; no client retries.
Quantization
format=NVFP4 W4A4 recipe=gittensor-model-hub RTX 5090 target checkpoint; BF16 speculative drafter is a runtime modification checkpoint=NVIDIA ModelOpt
Model revision
0cc27958cefbbe231782ec8511de8c4eb5233348
Evidence
Raw benchmark evidence attached · eligible and ranked
Published
by anonymous-EVjB2S

Benchmark command

sglang bench serving \
  --backend sglang-oai \
  --base-url http://127.0.0.1:9000 \
  --model gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --tokenizer gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --dataset-name custom \
  --dataset-path <EXACT_SYNTHETIC_DATASET> \
  --sharegpt-output-len 256 \
  --num-prompts 256 \
  --max-concurrency 1 \
  --request-rate inf \
  --seed 0 \
  --temperature 0 \
  --disable-tqdm \
  --output-file <local-path> \
  --warmup-requests 16 \
  --tag runpile-rtx5090-qwen38-dflash2-pilot-20260829-v2:qwen38-27b-nvfp4-sglang-dflash2-bf16-tp1:baseline:serve-fixed-1k-256-v1:c1:r1