gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

completed
Request throughput
3.52 req/s
Output throughput
3.5 tok/s
p95 TTFT
629.4 ms
p95 TPOT
0.00 ms

Manifest warnings

Review these published caveats before interpreting or comparing the metrics.

  • Review imported facts against the local model, hardware, launch command, and benchmark command before publication.

Server launch command

CUDA_VISIBLE_DEVICES=0 sglang serve \
  --model-path gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --revision 0cc27958cefbbe231782ec8511de8c4eb5233348 \
  --served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --port 9000 \
  --host 127.0.0.1 \
  --dtype bfloat16 \
  --mem-fraction-static 0.9 \
  --context-length 9216 \
  --max-running-requests 8 \
  --tp 1 \
  --kv-cache-dtype fp8_e4m3 \
  --chunked-prefill-size 1024 \
  --attention-backend flashinfer \
  --cuda-graph-max-bs 8 \
  --trust-remote-code \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-full-memory-ratio 5.61 \
  --mamba-radix-cache-strategy extra_buffer_lazy \
  --mamba-ssm-dtype bfloat16 \
  --max-mamba-cache-size 32 \
  --cuda-graph-bs-prefill 64 128 256 512 1024 \
  --enable-memory-saver \
  --enable-hierarchical-cache \
  --hicache-ratio 2 \
  --hicache-write-policy write_through \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal \
  --speculative-draft-model-revision b45a9e0fe2d62963a91158e6fb776fb86532c190 \
  --speculative-draft-model-quantization modelopt_fp4 \
  --speculative-num-draft-tokens 8

Run details

Workload
4096 → 1
Methodology
closed loop · streaming
Test load
2 concurrent requests
Endpoint
sglang-native-completions
Requests
16 unmeasured warmups per invocation after one readiness request.
Dataset
generator=sglang-benchmark-random-token-ids-fixed-v1 synthetic=true generator_revision=434a1959e7dc8814dfc5c6417830bb4b41aee21a608a0f578542ed9df76ea59d shared_prefix_tokens=0
Sampling
mode=greedy temperature=0
Stopping
ignore_eos=true max_output_tokens=1
Prefix and cache
cache_state=warm-host cache_enabled=false shared_prefix_tokens=0
Client
name=sglang benchmark serving backend=sglang version=0.0.0.dev1+g5f55db35e placement=same-host
Hardware
1× NVIDIA GeForce RTX 5090
Runtime
sglang 0.0.0.dev1+g5f55db35e
Decoding method
DFlash2 · calibrated NVFP4 draft · 8 draft tokens
Draft model
maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal · calibrated NVFP4revision b45a9e0fe2d62963a91158e6fb776fb86532c190
Runtime configuration
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Measurement
sglang benchmark serving 0.0.0.dev1+g5f55db35e
Timing
client-observed streaming request latency · client overhead included
Repetitions
1
Failure handling
Failed requests counted; no client retries.
Quantization
format=NVFP4 W4A4 recipe=gittensor-model-hub RTX 5090 target checkpoint; calibrated NVFP4 speculative drafter is a runtime modification checkpoint=NVIDIA ModelOpt
Model revision
0cc27958cefbbe231782ec8511de8c4eb5233348
Evidence
Raw benchmark evidence attached · eligible and ranked
Published
by anonymous-EVjB2S

Benchmark command

sglang bench serving \
  --backend sglang \
  --base-url http://127.0.0.1:9000 \
  --model gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --tokenizer gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
  --dataset-name random-ids \
  --dataset-path <EXACT_SYNTHETIC_DATASET> \
  --random-input-len 4096 \
  --random-output-len 1 \
  --random-range-ratio 1 \
  --tokenize-prompt \
  --num-prompts 256 \
  --max-concurrency 2 \
  --request-rate inf \
  --seed 0 \
  --temperature 0 \
  --disable-tqdm \
  --output-file <local-path> \
  --warmup-requests 16 \
  --tag runpile-rtx5090-qwen38-dflash2-concurrency-20260829-v5:qwen38-27b-nvfp4-sglang-dflash2-nvfp4-tp1:baseline:serve-fixed-4k-1-categorization-v1:c2:r2