A verifier that only ever passes or fails is a scorer. The third verdict — this instrument cannot decide, and here is why — is the one a motivated party would never build, so it is the one worth counting. Here is every refusal on record, by kind, each with the denominator it belongs to.
KINDS ARE NEVER MERGED on this page. REFUSED (our instrument declined) is not NEEDS DATA (the claimant published nothing to decide on) is not OPEN (a budget ran out, and the attempt is recorded) is not UNDECIDED (a swept cell that did not close) is not a scope limit (an audit that reached a fragment and says which). There is deliberately no total: one big number would be a bigger number and a smaller fact. Every count below is read from its record at this page's own build.
| kind | where | count | out of | record |
|---|---|---|---|---|
| REFUSED | the engine loop (keller-fibers) | 1 / 16,943 | objects certified this build | ledger.json |
| REFUSED | the matmul eval board | 39 / 301 | claims submitted | certs/matmul-eval-ledger.jsonl |
| NEEDS DATA | the kissing ledger, dimension 11 | 1 / 11 | ledger rows | certs/kissing-ledger.json |
| OPEN | the Erdős #1038 supremum campaign | 1 / 7 | degrees attempted | certs/sublevel-tao179.json |
| OPEN | the λ(6) campaign | 1 / 10 | families in the worklist | certs/lambda56-campaign.json |
| UNDECIDED | the two-population regime map | 9,915 / 21,567 | cells swept | certs/mfg2p-regime-map.json |
| UNDECIDED | the MFG regime map | 8,086 / 19,800 | cells swept | certs/mfg-regime-map.json |
| SCOPE | the six-claim AI audit | 6 / 6 | lanes audited | certs/ai-claims-summary.json |
Reading down the kind column is the point. These are five different behaviours that a single "refusals" counter would flatten into one: an instrument declining, a claimant publishing nothing, a budget running out, a sweep cell staying open, and an audit stating its reach. Only the first is a statement about the instrument.
The matmul eval board takes proposed exact rank-R decompositions from frontier models and decides each one in exact rational arithmetic. Of 364 real-model replies on record: 262 were decided (115 CERTIFIED, 147 REFUTED with the violated equation printed), and 39 were REFUSED because the reply carried no parseable proposal. That is a refusal rate of 13.0% on 301 submitted claims.
Two other outcomes sit in the same ledger and are NOT counted in that rate, because they are not our refusals: 23 replies in which the MODEL declined — it argued the target was impossible, which for the rank-6 ⟨2,2,2⟩ rung is correct and is Winograd's theorem — and 40 replies cut off by OUR OWN output cap, a harness artifact that says nothing about either party. Folding either of those into a refusal rate would inflate it with someone else's behaviour, or with our own plumbing.
A refusal rate with no denominator is unfalsifiable, and a refusal count with the wrong denominator is worse than none. Every row on this page names the population it is a fraction of, and the build refuses a row that cannot.
The sharpest refusal in the lab is not a verdict at all — it is the sentence that says which part of a claim was checked. All 6 lanes of the six-claim AI audit carry one, and 1 of them ends at PARTIAL. A reader who takes "CONFIRMED" off this table without its scope column has been told something the record does not say.
| claim | verdict | what was actually checked | named checks |
|---|---|---|---|
| Maxwell's point-charge bound | CONFIRMED | at ε = 1/6 only | 17 |
| The Korenblum constant | CONFIRMED | the numerical criterion | 15 |
| Erdős Problem #1038 | CONFIRMED | the computational fragment | 22 |
| Ran–Teng Conjecture 20 | PARTIAL | machine-checkable fragment only | 38 |
| The Mathieu property for Lie groups | CONFIRMED | supporting identities only | null |
| The rank-two Poisson conjecture | CONFIRMED | the explicit counterexample | null |
The two mean-field regime maps sweep 41,367 parameter cells and decide each one by an exhaustive box argument. 18,001 of them — 43.5% — did not close at the budget, and they are drawn on the atlas as holes rather than filled in. A map with no holes drawn is a map that interpolated somewhere, and the reader cannot tell where.
The same discipline applies to whole campaigns. The Erdős #1038 supremum ladder closed six degrees and RECORDS the one it attempted and could not finish (deg9 — bnb: box budget exhausted); the λ(6) campaign has 9 of its 10 families closed and publishes the remainder as unfinished rather than waiting for a clean number. An attempted-and-failed row is worth more than a silent gap: it tells the next person where the budget actually broke.