Comparison semantics · version 1
Comparability is computed, not assumed.
You may inspect any runs together. Runpile classifies material differences against your intent so a convenient table never becomes a misleading ranking.
Direct ranking is misleading
Examples: warm shared-prefix cache versus cold unique prompts; generation versus embedding.
Useful only with a prominent caveat
Examples: fixed 6,000/1 versus 6,144/1; synthetic versus real data; different timing boundaries.
Likely comparable with a caveat
Examples: a runtime patch-level change where workload and metric definitions match.
An intentional comparison dimension
Examples: RTX 5090 versus H100 when the selected intent is comparing hardware.