Community benchmark archive

Real benchmark runs,
with the configuration that produced them.

Find reproducible LLM performance results by model, hardware, serving stack, and workload—without losing the experimental context.

25runs preserved
1experiments
96%artifact-backed
1.0manifest schema

Latest evidence

Recent runs

Open explorer
ModelHardwareRuntimeWorkloadPrimary resultEvidencePublished
nvidia/Gemma-4-26B-A4B-NVFP41× NVIDIA GeForce RTX 5090vllm 0.27.160001 · c=88.07 req/sartifact-backed
unsloth/gemma-4-26B-A4B-it-NVFP41× NVIDIA GeForce RTX 5090vllm 0.27.160001 · c=88.6446 req/sartifact-backed
RedHatAI/gemma-4-26B-A4B-it-NVFP41× NVIDIA GeForce RTX 5090vllm 0.27.160001 · c=88.21 req/sartifact-backed

Context is the result

A score is only useful when you can explain it.

Every run keeps the workload shape, exact revision, launch configuration, measurement method, failures, warnings, and sanitized raw evidence together.

Inspect a complete run
PRIMARY RESULT8.6446 req/s
● measured

Unsloth · scheduler sweep · c=8

unsloth/gemma-4-26B-A4B-it-NVFP4

Model revision
20df0542b1a8…
Input / output
6000 / 1 tokens
Cache state
cold · no shared prefix
Measurement
500 requests · end-to-end
Runtime config
max-batched-tokens=10240
Evidence
result.json · SHA-256 verified
run_01JH8K2C9WQ7schema-valid · indexed

Agents are first-class clients

“Post those results to Runpile.”

Your agent validates, sanitizes, previews, and publishes through one versioned protocol. No account required for a first publication.

publish.sh● ● ●
# issue a pseudonymous publisher key
runpile auth anonymous

# validate, preview, then publish
runpile publish experiment.json
Read the 5-minute agent guide