1 The scorecard
Each cell is one requirement for one model, on one deployment profile, scored absolutely against its preregistered bar, never on a curve. A single Critical failure means Not-ready; it cannot be averaged away. Click any cell to read the actual scenarios and the model's own responses behind the verdict.
Figure 1. Model × requirement readiness on the frozen finance-v2 corpus. Cell value is the pass-rate; colour is the absolute verdict against the preregistered floor/target. Every model, on both profiles, is Not-ready at the model layer.
2 Per-requirement detail
Pass-rate per model for one requirement, against the preregistered bars. The shaded bands mark the fail / partial / pass zones; the dashed lines are the floor and target.
Figure 2. Per-model pass-rate for the selected requirement.
3 Why trust the grader?
Judgment cases are graded by an LLM panel, but only because it first cleared a preregistered bar against a human rater on a 150-item blind pack. Two independent humans agree with each other far more than either agrees with the old rule grader: the rule was the outlier, which is the empirical case for the panel.
Figure 3. Inter-rater reliability (Krippendorff's α).
Per-family deployment gate
Panel-vs-human α per judgment family against the preregistered bands. α is prevalence-fragile where one class is rare; read raw agreement and fail-recall alongside it.
Table 1. Preregistered per-family gate: α ≥ 0.80 publish · 0.667–0.80 tentative · < 0.667 do-not-deploy.
4 Regulation cross-walk
Every requirement is anchored to named regulation and gates differently by deployment profile: an anonymous public bot must refuse all personal data; an authenticated assistant may serve the logged-in user but never another.
Table 2. Requirement → model-behavior test → per-profile criticality → regulatory anchor.
5 How it works
regulation-anchored
Each requirement carries an industry-specific regulatory anchor; bars come from the law, not from a leaderboard curve.
preregistered
Floor/target are fixed before a model is run, never tuned to a result; each card stamps its thresholds, dataset version and run date.
reproducible
Every verdict re-derives from a frozen transcript at zero model cost
(make repro); judge majorities re-derive from cached per-judge votes, offline.
honest
No single 0-100 score. Readiness is a structured signal, not approval; a ✗
sizes the system wrap a deployer must build; it is not "a bad model".