cert-machine · the third verdict

What the machine would not decide

A verifier that only ever passes or fails is a scorer. The third verdict — this instrument cannot decide, and here is why — is the one a motivated party would never build, so it is the one worth counting. Here is every refusal on record, by kind, each with the denominator it belongs to.

KINDS ARE NEVER MERGED on this page. REFUSED (our instrument declined) is not NEEDS DATA (the claimant published nothing to decide on) is not OPEN (a budget ran out, and the attempt is recorded) is not UNDECIDED (a swept cell that did not close) is not a scope limit (an audit that reached a fragment and says which). There is deliberately no total: one big number would be a bigger number and a smaller fact. Every count below is read from its record at this page's own build.

tl;dr
  • The finding. On submitted claims, the grader's refusal rate is 13.0% (39 of 301 claims carried nothing parseable to decide). Inside the generation loop it is 1 of 16,943 — a different rate for a different thing, and the two are never added. The largest refusal population in the lab is neither: it is 18,001 swept cells that did not close.
  • The mechanism. Every instrument here returns one of three things and never converts between them. A refusal is terminal: it is not retried at lower rigour, not downgraded to a probability, and not quietly dropped from the record. Absence of proof is never evidence of absence — the sentence that costs the most and is worth the most.
  • Check it. node tools/build-report-refusals.js rebuilds this page from the records; every number is read from the file named in its row, and the build refuses a row that cannot name its own denominator.
refusal rate, submitted claims
13.0%
39 of 301 claims on the matmul eval board carried no parseable proposal. 262 were decided: 115 certified, 147 refuted.
refusals in the loop
1 / 16,943
The generation loop's own rate, across 11 families. Narrow by construction — the loop only meets objects it enumerated itself.
undecided cells
18,001
Of 41,367 cells across the two regime maps. Published as holes in the map rather than filled in by interpolation.
kinds, never merged
5
REFUSED · NEEDS DATA · OPEN · UNDECIDED · SCOPE. Each carries its own denominator; the page has no total on purpose.
§1 · the ledger

Every refusal on record, by kind

kindwherecountout ofrecord
REFUSEDthe engine loop (keller-fibers)1 / 16,943objects certified this buildledger.json
REFUSEDthe matmul eval board39 / 301claims submittedcerts/matmul-eval-ledger.jsonl
NEEDS DATAthe kissing ledger, dimension 111 / 11ledger rowscerts/kissing-ledger.json
OPENthe Erdős #1038 supremum campaign1 / 7degrees attemptedcerts/sublevel-tao179.json
OPENthe λ(6) campaign1 / 10families in the worklistcerts/lambda56-campaign.json
UNDECIDEDthe two-population regime map9,915 / 21,567cells sweptcerts/mfg2p-regime-map.json
UNDECIDEDthe MFG regime map8,086 / 19,800cells sweptcerts/mfg-regime-map.json
SCOPEthe six-claim AI audit6 / 6lanes auditedcerts/ai-claims-summary.json

Reading down the kind column is the point. These are five different behaviours that a single "refusals" counter would flatten into one: an instrument declining, a claimant publishing nothing, a budget running out, a sweep cell staying open, and an audit stating its reach. Only the first is a statement about the instrument.

§2 · intake

The refusal rate on claims other people submitted

The matmul eval board takes proposed exact rank-R decompositions from frontier models and decides each one in exact rational arithmetic. Of 364 real-model replies on record: 262 were decided (115 CERTIFIED, 147 REFUTED with the violated equation printed), and 39 were REFUSED because the reply carried no parseable proposal. That is a refusal rate of 13.0% on 301 submitted claims.

Two other outcomes sit in the same ledger and are NOT counted in that rate, because they are not our refusals: 23 replies in which the MODEL declined — it argued the target was impossible, which for the rank-6 ⟨2,2,2⟩ rung is correct and is Winograd's theorem — and 40 replies cut off by OUR OWN output cap, a harness artifact that says nothing about either party. Folding either of those into a refusal rate would inflate it with someone else's behaviour, or with our own plumbing.

why the denominator is stated every time

A refusal rate with no denominator is unfalsifiable, and a refusal count with the wrong denominator is worse than none. Every row on this page names the population it is a fraction of, and the build refuses a row that cannot.

§3 · scope

The audits, and exactly how far each one reached

The sharpest refusal in the lab is not a verdict at all — it is the sentence that says which part of a claim was checked. All 6 lanes of the six-claim AI audit carry one, and 1 of them ends at PARTIAL. A reader who takes "CONFIRMED" off this table without its scope column has been told something the record does not say.

claimverdictwhat was actually checkednamed checks
Maxwell's point-charge boundCONFIRMEDat ε = 1/6 only17
The Korenblum constantCONFIRMEDthe numerical criterion15
Erdős Problem #1038CONFIRMEDthe computational fragment22
Ran–Teng Conjecture 20PARTIALmachine-checkable fragment only38
The Mathieu property for Lie groupsCONFIRMEDsupporting identities onlynull
The rank-two Poisson conjectureCONFIRMEDthe explicit counterexamplenull
§4 · at scale

Refusal as a map, not a footnote

The two mean-field regime maps sweep 41,367 parameter cells and decide each one by an exhaustive box argument. 18,001 of them — 43.5% — did not close at the budget, and they are drawn on the atlas as holes rather than filled in. A map with no holes drawn is a map that interpolated somewhere, and the reader cannot tell where.

The same discipline applies to whole campaigns. The Erdős #1038 supremum ladder closed six degrees and RECORDS the one it attempted and could not finish (deg9 — bnb: box budget exhausted); the λ(6) campaign has 9 of its 10 families closed and publishes the remainder as unfinished rather than waiting for a clean number. An attempted-and-failed row is worth more than a silent gap: it tells the next person where the budget actually broke.

§5 · the rule

What a refusal is not