interruptedrun_g4_024published 8/16/2026
unsloth/gemma-4-26B-A4B-it-NVFP4
Interrupted preliminary screening
4× NVIDIA GeForce RTX 5090 vllm 0.27.1 6000 in → 1 out c=8
Primary result—No performance result
35 evidence score
01 / Measurement
Results
Submitted metrics retain their scope, statistic, unit, and measurement status.
Run outcome—interrupted
02 / Experimental shape
Workload
The work performed is part of the result, not footnote metadata.
- Mode
- online
- Task
- text_generation
- Dataset
- {"license":"CC0-1.0","generator":"vllm-random","synthetic":true}
- Sampling
- {"temperature":0,"max_output_tokens":1}
- Streaming
- no
- Concurrency
- 8
- Random Seed
- 42
- Input Tokens
- {"kind":"fixed","value":6000}
- Output Tokens
- {"kind":"fixed","value":1}
- Request Count
- 500
- Arrival Process
- closed_loop
- Prefix Behavior
- {"cache_state":"cold","cache_enabled":false,"shared_prefix_tokens":0}
- Warmup Request Count
- 8
- Request Rate Per Second
- unknown
03 / Serving stack
Runtime
Exact versions and namespaced configuration remain available for reproduction.
- Engine
- vllm
- Version
- 0.27.1
- Parallelism
- unknown
- Parameters
- {"vllm":{"max_num_batched_tokens":8192}}
04 / Environment
Hardware
Host and accelerator context associated with this execution.
- Count
- 4
- Model
- GeForce RTX 5090
- Driver
- 575.64.03
- Vendor
- NVIDIA
- Architecture
- Blackwell
- Interconnect
- PCIe 5.0; no NVLink
- Device Indices
- [0,1,2,3]
- Power Limit Watts
- 575
- Memory Bytes Per Device
- 34359738368
- Os
- Ubuntu 24.04.2 LTS
- Cpu
- {"model":"AMD Ryzen Threadripper PRO","logical_cores":64,"physical_cores":32}
- Kernel
- 6.11.0-26-generic
- Topology
- Four discrete PCIe GPUs, one serving process per GPU.
- Driver Versions
- {"nvidia":"575.64.03","cuda_toolkit":"12.9"}
- Host Architecture
- x86_64
- System Memory Bytes
- 274877906944
- Virtual Environment
- {"python":"3.12.10","dependency_lock_hash":"074c2809a1a0ddf856a6b751e0b20773f8cc0828e7015e49d4daa7678dc38f54"}
05 / Reproduction
Command
Sanitized before publication. Local paths, hosts, and credentials are omitted.
No sanitized launch command was supplied.
Measurement method
{
"harness": "vllm bench serve",
"version": "0.27.1",
"clock_source": "CLOCK_MONOTONIC",
"warmup_semantics": "Eight requests completed before the timed batch.",
"duration_semantics": "First timed request dispatch through final response completion.",
"tokenization_included": true,
"model_loading_included": false,
"retry_failure_treatment": "Failed requests remain in the failure count and are not retried.",
"client_overhead_included": true,
"metric_definition_version": "runpile-generation-1"
}06 / Provenance
Evidence
Artifact display names are descriptive; hashes and Runpile IDs are authoritative.
unavailablescreening-interrupted.jsonbenchmark_result_json · 18,432 bytes
01541ded8566c149dd63e6cd577bef82790752293bc07cbc056e457c9c101686demo evidenceComplete submitted run JSON
{
"name": "Interrupted preliminary screening",
"status": "interrupted",
"runtime": {
"engine": "vllm",
"version": "0.27.1",
"parameters": {
"vllm": {
"max_num_batched_tokens": 8192
}
}
},
"subject": {
"type": "model",
"source": "huggingface",
"revision": "20df0542b1a86ce19f495ac2eca2c7c12bce82f9",
"identifier": "unsloth/gemma-4-26B-A4B-it-NVFP4",
"revision_kind": "commit"
},
"relation": {
"type": "other",
"group": "screening"
},
"warnings": [
"Interrupted after 73 of 500 requests; excluded from completed-run rankings."
],
"workload": {
"mode": "online",
"task": "text_generation",
"concurrency": 8,
"input_tokens": {
"kind": "fixed",
"value": 6000
},
"output_tokens": {
"kind": "fixed",
"value": 1
},
"request_count": 500
},
"artifacts": [
{
"kind": "benchmark_result_json",
"sha256": "01541ded8566c149dd63e6cd577bef82790752293bc07cbc056e457c9c101686",
"license": "CC-BY-4.0",
"filename": "screening-interrupted.json",
"media_type": "application/json",
"size_bytes": 18432,
"visibility": "public",
"run_client_id": "screening-interrupted",
"client_artifact_id": "screening-interrupted-raw",
"raw_requests_included": false
}
],
"client_run_id": "screening-interrupted"
}