- Request throughput
- 0.83 req/s
- Output throughput
- 854.3 tok/s
- p95 TTFT
- 58.1 ms
- p95 TPOT
- 4.06 ms
Manifest warnings
Review these published caveats before interpreting or comparing the metrics.
- Review imported facts against the local model, hardware, launch command, and benchmark command before publication.
Server launch command
CUDA_VISIBLE_DEVICES=0 sglang serve \
--model-path gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--revision 0cc27958cefbbe231782ec8511de8c4eb5233348 \
--served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--port 9000 \
--host 127.0.0.1 \
--dtype bfloat16 \
--mem-fraction-static 0.9 \
--context-length 9216 \
--max-running-requests 8 \
--tp 1 \
--kv-cache-dtype fp8_e4m3 \
--chunked-prefill-size 1024 \
--attention-backend flashinfer \
--cuda-graph-max-bs 8 \
--trust-remote-code \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-full-memory-ratio 5.61 \
--mamba-radix-cache-strategy extra_buffer_lazy \
--mamba-ssm-dtype bfloat16 \
--max-mamba-cache-size 32 \
--cuda-graph-bs-prefill 64 128 256 512 1024 \
--enable-memory-saver \
--enable-hierarchical-cache \
--hicache-ratio 2 \
--hicache-write-policy write_through \
--speculative-algorithm DFLASH \
--speculative-draft-model-path maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal \
--speculative-draft-model-revision b45a9e0fe2d62963a91158e6fb776fb86532c190 \
--speculative-draft-model-quantization modelopt_fp4 \
--speculative-num-draft-tokens 8Run details
- Workload
- 256 → 1024
- Methodology
- closed loop · streaming
- Test load
- 2 concurrent requests
- Endpoint
openai-compatible-completions- Requests
- 8 unmeasured warmups per invocation after one readiness request.
- Dataset
sha256=20954ec817c4b70aa26473fd6b50c397ddc8acd8ccd2e58b1c74128586b39b62 generator=runpile-sweeps-tokenizer-verified-jsonl synthetic=true generator_revision=b62df3a992da3d4a1dc6f2f12ce2706bbeef61bc5ca0c03c917d7597606c3312 shared_prefix_tokens=0- Sampling
mode=greedy temperature=0- Stopping
ignore_eos=true max_output_tokens=1024- Prefix and cache
cache_state=warm-host cache_enabled=false shared_prefix_tokens=0- Client
name=sglang benchmark serving backend=sglang-oai version=0.0.0.dev1+g5f55db35e placement=same-host- Hardware
- 1× NVIDIA GeForce RTX 5090
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Decoding method
- DFlash2 · calibrated NVFP4 draft · 8 draft tokens
- Draft model
- maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal · calibrated NVFP4revision
b45a9e0fe2d62963a91158e6fb776fb86532c190 - Runtime configuration
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Measurement
- sglang benchmark serving 0.0.0.dev1+g5f55db35e
- Timing
- client-observed streaming request latency · client overhead included
- Repetitions
- 1
- Failure handling
- Failed requests counted; no client retries.
- Quantization
format=NVFP4 W4A4 recipe=gittensor-model-hub RTX 5090 target checkpoint; calibrated NVFP4 speculative drafter is a runtime modification checkpoint=NVIDIA ModelOpt- Model revision
0cc27958cefbbe231782ec8511de8c4eb5233348- Evidence
- Raw benchmark evidence attached · eligible and ranked
- Published
- by anonymous-EVjB2S
Benchmark command
sglang bench serving \
--backend sglang-oai \
--base-url http://127.0.0.1:9000 \
--model gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--tokenizer gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--dataset-name custom \
--dataset-path <EXACT_SYNTHETIC_DATASET> \
--sharegpt-output-len 1024 \
--num-prompts 128 \
--max-concurrency 2 \
--request-rate inf \
--seed 0 \
--temperature 0 \
--disable-tqdm \
--output-file <local-path> \
--warmup-requests 8 \
--tag runpile-rtx5090-qwen38-dflash2-concurrency-20260829-v5:qwen38-27b-nvfp4-sglang-dflash2-nvfp4-tp1:baseline:serve-fixed-256-1k-v1:c2:r1