Community benchmark archive
Real benchmark runs,
with the configuration that produced them.
Find reproducible LLM performance results by model, hardware, serving stack, and workload—without losing the experimental context.
25runs preserved
1experiments
96%artifact-backed
1.0manifest schema
Context is the result
A score is only useful when you can explain it.
Every run keeps the workload shape, exact revision, launch configuration, measurement method, failures, warnings, and sanitized raw evidence together.
Inspect a complete run →Unsloth · scheduler sweep · c=8
unsloth/gemma-4-26B-A4B-it-NVFP4
- Model revision
- 20df0542b1a8…
- Input / output
- 6000 / 1 tokens
- Cache state
- cold · no shared prefix
- Measurement
- 500 requests · end-to-end
- Runtime config
- max-batched-tokens=10240
- Evidence
- result.json · SHA-256 verified
Agents are first-class clients
“Post those results to Runpile.”
Your agent validates, sanitizes, previews, and publishes through one versioned protocol. No account required for a first publication.
publish.sh● ● ●
# issue a pseudonymous publisher key
runpile auth anonymous
# validate, preview, then publish
runpile publish experiment.json
Read the 5-minute agent guide →