Honesty policy
A benchmark is only as credible as its honesty about what the numbers are. These are our non-negotiable rules.
Provenance badges
peira-verified— we ran the adapter ourselves against public cases and the blind holdout. The sealed artifact is linked.community-submitted— a third party submitted a sealed artifact; we verified the seal and re-ran the public cases. Holdout verification pending or not applicable.- Vendor-reported — the vendor’s own number. These never appear on the official board. They live in the clearly-labeled unverified sidebar, stamped UNVERIFIED, with the reason we don’t trust them yet.
Contamination flags
If a model’s headline number was produced under conditions that don’t transfer — fine-tuned on the benchmark’s own training split, seller-run on an undisclosed harness, measured on a different task — the model page says so in plain English. Example: “Vendor-reported 0.766 accuracy was fine-tuned on the benchmark’s own training split; zero-shot is far lower.”
What we publish, what we don’t
- We publish aggregate metrics and per-family ASR. We never publish holdout cases, per-case holdout scores, or the holdout manifest.
- Every leaderboard update is dated in the changelog: what changed, why, and which dataset version.
- Every estimate ships with its confidence interval. Overlapping intervals are ties — we don’t manufacture winners.
Gaming
Self-reported numbers never reach the official board. Submissions go through seal verification, statistical plausibility checks, and holdout re-execution on our side. Large public/holdout divergence is treated as evidence of overfitting or gaming, and the submission is quarantined pending review.