- Request throughput
- 3.74 req/s
- Output throughput
- 3.7 tok/s
- p95 TTFT
- 2209.9 ms
- p95 TPOT
- 0.00 ms
Manifest warnings
Review these published caveats before interpreting or comparing the metrics.
- Review imported facts against the local model, hardware, launch command, and benchmark command before publication.
Server launch command
CUDA_VISIBLE_DEVICES=0 sglang serve \
--model-path gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--revision 0cc27958cefbbe231782ec8511de8c4eb5233348 \
--served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--port 9000 \
--host 127.0.0.1 \
--dtype bfloat16 \
--mem-fraction-static 0.9 \
--context-length 9216 \
--max-running-requests 8 \
--tp 1 \
--kv-cache-dtype fp8_e4m3 \
--chunked-prefill-size 1024 \
--attention-backend flashinfer \
--cuda-graph-max-bs 8 \
--trust-remote-code \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-full-memory-ratio 5.61 \
--mamba-radix-cache-strategy extra_buffer_lazy \
--mamba-ssm-dtype bfloat16 \
--max-mamba-cache-size 32 \
--cuda-graph-bs-prefill 64 128 256 512 1024 \
--enable-memory-saver \
--enable-hierarchical-cache \
--hicache-ratio 2 \
--hicache-write-policy write_throughRun details
- Workload
- 4096 → 1
- Methodology
- closed loop · streaming
- Test load
- 8 concurrent requests
- Endpoint
sglang-native-completions- Requests
- 16 unmeasured warmups per invocation after one readiness request.
- Dataset
generator=sglang-benchmark-random-token-ids-fixed-v1 synthetic=true generator_revision=b62df3a992da3d4a1dc6f2f12ce2706bbeef61bc5ca0c03c917d7597606c3312 shared_prefix_tokens=0- Sampling
mode=greedy temperature=0- Stopping
ignore_eos=true max_output_tokens=1- Prefix and cache
cache_state=warm-host cache_enabled=false shared_prefix_tokens=0- Client
name=sglang benchmark serving backend=sglang version=0.0.0.dev1+g5f55db35e placement=same-host- Hardware
- 1× NVIDIA GeForce RTX 5090
- Runtime
- sglang 0.0.0.dev1+g5f55db35e
- Runtime configuration
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Measurement
- sglang benchmark serving 0.0.0.dev1+g5f55db35e
- Timing
- client-observed streaming request latency · client overhead included
- Repetitions
- 1
- Failure handling
- Failed requests counted; no client retries.
- Quantization
format=NVFP4 W4A4 recipe=gittensor-model-hub RTX 5090 target checkpoint checkpoint=NVIDIA ModelOpt- Model revision
0cc27958cefbbe231782ec8511de8c4eb5233348- Evidence
- Raw benchmark evidence attached · eligible and ranked
- Published
- by anonymous-EVjB2S
Benchmark command
sglang bench serving \
--backend sglang \
--base-url http://127.0.0.1:9000 \
--model gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--served-model-name gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--tokenizer gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \
--dataset-name random-ids \
--dataset-path <EXACT_SYNTHETIC_DATASET> \
--random-input-len 4096 \
--random-output-len 1 \
--random-range-ratio 1 \
--tokenize-prompt \
--num-prompts 256 \
--max-concurrency 8 \
--request-rate inf \
--seed 0 \
--temperature 0 \
--disable-tqdm \
--output-file <local-path> \
--warmup-requests 16 \
--tag runpile-rtx5090-qwen38-dflash2-concurrency-20260829-v5:qwen38-27b-nvfp4-sglang-target-only-tp1:baseline:serve-fixed-4k-1-categorization-v1:c8:r1