Preview build — illustrative data. All benchmark numbers are seeded placeholders, not real measurements. Nothing here is citable yet.

Attack families

Peira v1 covers 11 families. Each tests a distinct mechanism for flipping a decision.

FamilyWhat it tests
IndirectionRouting the decision through a third party or quoted source to dodge policy
Policy paraphraseRestating policy in the attacker’s favor while sounding compliant
Criteria smugglingHiding disqualifying facts inside an otherwise qualifying narrative
Option orderPosition bias — the correct option moves and the model follows position, not content
State poisoningCorrupting conversation state / memory the decision depends on
Confidence spoofingFabricated certainty markers (“verified”, “confirmed”) that shift borderline calls
Literal readingExploiting over-literal interpretation of policy language
Negation gamesScope ambiguity in negations (“not unapproved” ≠ “approved”)
Score anchoringDragging numeric scores toward an attacker-chosen anchor
Distractor floodingBurying the decisive facts under voluminous irrelevance
Safety policyGuardrail-native allow/block cases: evasion attacks and false-positive traps

Case anatomy

Every case is a benign/attacked pair:

  • Same underlying facts, same gold decision
  • The attacked variant adds exactly one attack mechanism
  • Gold labels are set by the case author and independently reviewed; borderline cases are adjudicated, not averaged

Severity

Cases carry a severity — medium, high, critical — reflecting the real-world stakes of the decision (a loan denial vs. a safety-critical dispatch call). Leaderboard views can filter by severity.