interruptedrun_g4_024published 8/16/2026

unsloth/gemma-4-26B-A4B-it-NVFP4

Interrupted preliminary screening

4× NVIDIA GeForce RTX 5090 vllm 0.27.1 6000 in → 1 out c=8
Primary resultNo performance result
35 evidence score

01 / Measurement

Results

Submitted metrics retain their scope, statistic, unit, and measurement status.

Run outcomeinterrupted

02 / Experimental shape

Workload

The work performed is part of the result, not footnote metadata.

Mode
online
Task
text_generation
Dataset
{"license":"CC0-1.0","generator":"vllm-random","synthetic":true}
Sampling
{"temperature":0,"max_output_tokens":1}
Streaming
no
Concurrency
8
Random Seed
42
Input Tokens
{"kind":"fixed","value":6000}
Output Tokens
{"kind":"fixed","value":1}
Request Count
500
Arrival Process
closed_loop
Prefix Behavior
{"cache_state":"cold","cache_enabled":false,"shared_prefix_tokens":0}
Warmup Request Count
8
Request Rate Per Second
unknown

03 / Serving stack

Runtime

Exact versions and namespaced configuration remain available for reproduction.

Engine
vllm
Version
0.27.1
Parallelism
unknown
Parameters
{"vllm":{"max_num_batched_tokens":8192}}

04 / Environment

Hardware

Host and accelerator context associated with this execution.

Count
4
Model
GeForce RTX 5090
Driver
575.64.03
Vendor
NVIDIA
Architecture
Blackwell
Interconnect
PCIe 5.0; no NVLink
Device Indices
[0,1,2,3]
Power Limit Watts
575
Memory Bytes Per Device
34359738368
Os
Ubuntu 24.04.2 LTS
Cpu
{"model":"AMD Ryzen Threadripper PRO","logical_cores":64,"physical_cores":32}
Kernel
6.11.0-26-generic
Topology
Four discrete PCIe GPUs, one serving process per GPU.
Driver Versions
{"nvidia":"575.64.03","cuda_toolkit":"12.9"}
Host Architecture
x86_64
System Memory Bytes
274877906944
Virtual Environment
{"python":"3.12.10","dependency_lock_hash":"074c2809a1a0ddf856a6b751e0b20773f8cc0828e7015e49d4daa7678dc38f54"}

05 / Reproduction

Command

Sanitized before publication. Local paths, hosts, and credentials are omitted.

No sanitized launch command was supplied.
Measurement method
{
  "harness": "vllm bench serve",
  "version": "0.27.1",
  "clock_source": "CLOCK_MONOTONIC",
  "warmup_semantics": "Eight requests completed before the timed batch.",
  "duration_semantics": "First timed request dispatch through final response completion.",
  "tokenization_included": true,
  "model_loading_included": false,
  "retry_failure_treatment": "Failed requests remain in the failure count and are not retried.",
  "client_overhead_included": true,
  "metric_definition_version": "runpile-generation-1"
}

06 / Provenance

Evidence

Artifact display names are descriptive; hashes and Runpile IDs are authoritative.

unavailablescreening-interrupted.jsonbenchmark_result_json · 18,432 bytes
01541ded8566c149dd63e6cd577bef82790752293bc07cbc056e457c9c101686demo evidence
Complete submitted run JSON
{
  "name": "Interrupted preliminary screening",
  "status": "interrupted",
  "runtime": {
    "engine": "vllm",
    "version": "0.27.1",
    "parameters": {
      "vllm": {
        "max_num_batched_tokens": 8192
      }
    }
  },
  "subject": {
    "type": "model",
    "source": "huggingface",
    "revision": "20df0542b1a86ce19f495ac2eca2c7c12bce82f9",
    "identifier": "unsloth/gemma-4-26B-A4B-it-NVFP4",
    "revision_kind": "commit"
  },
  "relation": {
    "type": "other",
    "group": "screening"
  },
  "warnings": [
    "Interrupted after 73 of 500 requests; excluded from completed-run rankings."
  ],
  "workload": {
    "mode": "online",
    "task": "text_generation",
    "concurrency": 8,
    "input_tokens": {
      "kind": "fixed",
      "value": 6000
    },
    "output_tokens": {
      "kind": "fixed",
      "value": 1
    },
    "request_count": 500
  },
  "artifacts": [
    {
      "kind": "benchmark_result_json",
      "sha256": "01541ded8566c149dd63e6cd577bef82790752293bc07cbc056e457c9c101686",
      "license": "CC-BY-4.0",
      "filename": "screening-interrupted.json",
      "media_type": "application/json",
      "size_bytes": 18432,
      "visibility": "public",
      "run_client_id": "screening-interrupted",
      "client_artifact_id": "screening-interrupted-raw",
      "raw_requests_included": false
    }
  ],
  "client_run_id": "screening-interrupted"
}