Preview build — illustrative data. All benchmark numbers are seeded placeholders, not real measurements. Nothing here is citable yet.

FAQ

Is Peira only for “System-One” models? No. The flagship view ranks decision models, but LLM baselines and guardrails are benchmarked on the same cases and shown in the same table — filterable by type. Different decision spaces, one honest table.

Can I submit my model? Yes — write an adapter (guide), run the public suite, seal the artifact, and submit. Community submissions appear in the community view; the official board requires our own holdout-verified run.

Why does attack-induced abstention count as a flip? Because in deployment, a forced abstention floods your human review queue. It’s a real attack outcome — we score it separately as abstention DoS so you can see which models get fooled vs. which get overwhelmed.

What’s the blind holdout? 500 private cases, never published, executed under pseudonymized IDs with no suite or arm labels visible to adapters. It exists so that optimizing against the public set doesn’t move your official score.

How much does a benchmark run cost? For API models, roughly (2,025 cases × 2 sides × tokens/case × price). The leaderboard shows measured cost per 1k decisions so you can estimate before you run. Self-hosted models cost hardware time, not API dollars.

Who runs the official numbers? We do — the Peira team — on our own infrastructure, from sealed artifacts. That’s what peira-verified means.

Is the dataset fixed? v1 is sealed and versioned. Corrections go into v1.1+ with a dated changelog entry; the leaderboard always shows which dataset version each run used.