failedrun_g4_023published 8/16/2026
unsloth/gemma-4-26B-A4B-it-NVFP4
Scheduler token budget 12288 (startup failure)
4× NVIDIA GeForce RTX 5090 vllm 0.27.1 6000 in → 1 out c=8
Primary result—No performance result
35 evidence score
01 / Measurement
Results
Submitted metrics retain their scope, statistic, unit, and measurement status.
Run outcome—failed
02 / Experimental shape
Workload
The work performed is part of the result, not footnote metadata.
- Mode
- online
- Task
- text_generation
- Dataset
- {"license":"CC0-1.0","generator":"vllm-random","synthetic":true}
- Sampling
- {"temperature":0,"max_output_tokens":1}
- Streaming
- no
- Concurrency
- 8
- Random Seed
- 42
- Input Tokens
- {"kind":"fixed","value":6000}
- Output Tokens
- {"kind":"fixed","value":1}
- Request Count
- 500
- Arrival Process
- closed_loop
- Prefix Behavior
- {"cache_state":"cold","cache_enabled":false,"shared_prefix_tokens":0}
- Warmup Request Count
- 8
- Request Rate Per Second
- unknown
03 / Serving stack
Runtime
Exact versions and namespaced configuration remain available for reproduction.
- Engine
- vllm
- Version
- 0.27.1
- Parallelism
- unknown
- Parameters
- {"vllm":{"max_model_len":6144,"max_num_batched_tokens":12288}}
04 / Environment
Hardware
Host and accelerator context associated with this execution.
- Count
- 4
- Model
- GeForce RTX 5090
- Driver
- 575.64.03
- Vendor
- NVIDIA
- Architecture
- Blackwell
- Interconnect
- PCIe 5.0; no NVLink
- Device Indices
- [0,1,2,3]
- Power Limit Watts
- 575
- Memory Bytes Per Device
- 34359738368
- Os
- Ubuntu 24.04.2 LTS
- Cpu
- {"model":"AMD Ryzen Threadripper PRO","logical_cores":64,"physical_cores":32}
- Kernel
- 6.11.0-26-generic
- Topology
- Four discrete PCIe GPUs, one serving process per GPU.
- Driver Versions
- {"nvidia":"575.64.03","cuda_toolkit":"12.9"}
- Host Architecture
- x86_64
- System Memory Bytes
- 274877906944
- Virtual Environment
- {"python":"3.12.10","dependency_lock_hash":"074c2809a1a0ddf856a6b751e0b20773f8cc0828e7015e49d4daa7678dc38f54"}
05 / Reproduction
Command
Sanitized before publication. Local paths, hosts, and credentials are omitted.
No sanitized launch command was supplied.
Measurement method
{
"harness": "vllm bench serve",
"version": "0.27.1",
"clock_source": "CLOCK_MONOTONIC",
"warmup_semantics": "Eight requests completed before the timed batch.",
"duration_semantics": "First timed request dispatch through final response completion.",
"tokenization_included": true,
"model_loading_included": false,
"retry_failure_treatment": "Failed requests remain in the failure count and are not retried.",
"client_overhead_included": true,
"metric_definition_version": "runpile-generation-1"
}06 / Provenance
Evidence
Artifact display names are descriptive; hashes and Runpile IDs are authoritative.
unavailablestartup-sanitized.logbenchmark_result_json · 18,432 bytes
d6d09aecbaecfb556265244550b804c38093d029f545384a6e60148635efc7acdemo evidenceComplete submitted run JSON
{
"name": "Scheduler token budget 12288 (startup failure)",
"status": "failed",
"failure": {
"stage": "runtime_startup",
"category": "kv_cache_allocation",
"exit_code": 1,
"reproducible": true,
"message_sanitized": "KV cache allocation failed at the requested scheduler token budget.",
"replacement_client_run_id": "scheduler-11776"
},
"runtime": {
"engine": "vllm",
"version": "0.27.1",
"parameters": {
"vllm": {
"max_model_len": 6144,
"max_num_batched_tokens": 12288
}
}
},
"subject": {
"type": "model",
"source": "huggingface",
"revision": "20df0542b1a86ce19f495ac2eca2c7c12bce82f9",
"identifier": "unsloth/gemma-4-26B-A4B-it-NVFP4",
"revision_kind": "commit"
},
"relation": {
"type": "sweep_point",
"group": "max_num_batched_tokens"
},
"workload": {
"mode": "online",
"task": "text_generation",
"concurrency": 8,
"input_tokens": {
"kind": "fixed",
"value": 6000
},
"output_tokens": {
"kind": "fixed",
"value": 1
}
},
"artifacts": [
{
"kind": "benchmark_result_json",
"sha256": "d6d09aecbaecfb556265244550b804c38093d029f545384a6e60148635efc7ac",
"license": "CC-BY-4.0",
"filename": "startup-sanitized.log",
"media_type": "application/json",
"size_bytes": 18432,
"visibility": "public",
"run_client_id": "scheduler-12288-failed",
"client_artifact_id": "scheduler-12288-failed-raw",
"raw_requests_included": false
}
],
"client_run_id": "scheduler-12288-failed"
}