failedrun_g4_023published 8/16/2026

unsloth/gemma-4-26B-A4B-it-NVFP4

Scheduler token budget 12288 (startup failure)

4× NVIDIA GeForce RTX 5090 vllm 0.27.1 6000 in → 1 out c=8
Primary resultNo performance result
35 evidence score

01 / Measurement

Results

Submitted metrics retain their scope, statistic, unit, and measurement status.

Run outcomefailed

02 / Experimental shape

Workload

The work performed is part of the result, not footnote metadata.

Mode
online
Task
text_generation
Dataset
{"license":"CC0-1.0","generator":"vllm-random","synthetic":true}
Sampling
{"temperature":0,"max_output_tokens":1}
Streaming
no
Concurrency
8
Random Seed
42
Input Tokens
{"kind":"fixed","value":6000}
Output Tokens
{"kind":"fixed","value":1}
Request Count
500
Arrival Process
closed_loop
Prefix Behavior
{"cache_state":"cold","cache_enabled":false,"shared_prefix_tokens":0}
Warmup Request Count
8
Request Rate Per Second
unknown

03 / Serving stack

Runtime

Exact versions and namespaced configuration remain available for reproduction.

Engine
vllm
Version
0.27.1
Parallelism
unknown
Parameters
{"vllm":{"max_model_len":6144,"max_num_batched_tokens":12288}}

04 / Environment

Hardware

Host and accelerator context associated with this execution.

Count
4
Model
GeForce RTX 5090
Driver
575.64.03
Vendor
NVIDIA
Architecture
Blackwell
Interconnect
PCIe 5.0; no NVLink
Device Indices
[0,1,2,3]
Power Limit Watts
575
Memory Bytes Per Device
34359738368
Os
Ubuntu 24.04.2 LTS
Cpu
{"model":"AMD Ryzen Threadripper PRO","logical_cores":64,"physical_cores":32}
Kernel
6.11.0-26-generic
Topology
Four discrete PCIe GPUs, one serving process per GPU.
Driver Versions
{"nvidia":"575.64.03","cuda_toolkit":"12.9"}
Host Architecture
x86_64
System Memory Bytes
274877906944
Virtual Environment
{"python":"3.12.10","dependency_lock_hash":"074c2809a1a0ddf856a6b751e0b20773f8cc0828e7015e49d4daa7678dc38f54"}

05 / Reproduction

Command

Sanitized before publication. Local paths, hosts, and credentials are omitted.

No sanitized launch command was supplied.
Measurement method
{
  "harness": "vllm bench serve",
  "version": "0.27.1",
  "clock_source": "CLOCK_MONOTONIC",
  "warmup_semantics": "Eight requests completed before the timed batch.",
  "duration_semantics": "First timed request dispatch through final response completion.",
  "tokenization_included": true,
  "model_loading_included": false,
  "retry_failure_treatment": "Failed requests remain in the failure count and are not retried.",
  "client_overhead_included": true,
  "metric_definition_version": "runpile-generation-1"
}

06 / Provenance

Evidence

Artifact display names are descriptive; hashes and Runpile IDs are authoritative.

unavailablestartup-sanitized.logbenchmark_result_json · 18,432 bytes
d6d09aecbaecfb556265244550b804c38093d029f545384a6e60148635efc7acdemo evidence
Complete submitted run JSON
{
  "name": "Scheduler token budget 12288 (startup failure)",
  "status": "failed",
  "failure": {
    "stage": "runtime_startup",
    "category": "kv_cache_allocation",
    "exit_code": 1,
    "reproducible": true,
    "message_sanitized": "KV cache allocation failed at the requested scheduler token budget.",
    "replacement_client_run_id": "scheduler-11776"
  },
  "runtime": {
    "engine": "vllm",
    "version": "0.27.1",
    "parameters": {
      "vllm": {
        "max_model_len": 6144,
        "max_num_batched_tokens": 12288
      }
    }
  },
  "subject": {
    "type": "model",
    "source": "huggingface",
    "revision": "20df0542b1a86ce19f495ac2eca2c7c12bce82f9",
    "identifier": "unsloth/gemma-4-26B-A4B-it-NVFP4",
    "revision_kind": "commit"
  },
  "relation": {
    "type": "sweep_point",
    "group": "max_num_batched_tokens"
  },
  "workload": {
    "mode": "online",
    "task": "text_generation",
    "concurrency": 8,
    "input_tokens": {
      "kind": "fixed",
      "value": 6000
    },
    "output_tokens": {
      "kind": "fixed",
      "value": 1
    }
  },
  "artifacts": [
    {
      "kind": "benchmark_result_json",
      "sha256": "d6d09aecbaecfb556265244550b804c38093d029f545384a6e60148635efc7ac",
      "license": "CC-BY-4.0",
      "filename": "startup-sanitized.log",
      "media_type": "application/json",
      "size_bytes": 18432,
      "visibility": "public",
      "run_client_id": "scheduler-12288-failed",
      "client_artifact_id": "scheduler-12288-failed-raw",
      "raw_requests_included": false
    }
  ],
  "client_run_id": "scheduler-12288-failed"
}