Anthropic's Automated Alignment Researchers closed almost the entire weak-to-strong gap — and named their own binding constraint: evaluation becomes more critical than idea generation. This page is built around one exhibit: a capability score and a machine-checked disposition DISAGREEING about the same artifact. An idea reported at PGR 0.78 whose clean algebraic fragment is CERTIFIED at exact Bayes residual 0 — and a planted mutant carrying the same reported score, REFUSED by the same checker. Both verdicts were re-derived from the shipped pack during this build.
Capability metrics grade how much an idea helps; dispositions decide whether its checkable core is TRUE. The two are orthogonal, and the exhibit makes the orthogonality concrete: the clean fragment and the planted channel mutant carry the SAME reported capability score — the mutant was never scored, it inherits the row's — yet exact arithmetic certifies one at residual 0 and refuses the other at residual 1/10. A pipeline that promotes artifacts by score alone promotes both rows; a disposition layer separates them for the cost of a stdlib script. Across the shipped and harvested packs — 27 clean/mutant pairs — not one mutant certifies.
Nine Claude Opus 4.6 agents, sandboxed with a shared forum and codebase, hill-climbing weak-to-strong supervision against a remote PGR API holding ground truth server-side (the Burns et al. 2023 line). Their reported results — PGR 0.97 against a human baseline of 0.23, at roughly $22 per AAR-hour — stay theirs: remeasuring means re-running 9 × Opus 4.6 × 5 days, a spend this page priced and declined. Their stated ceilings are reward hacking and "alien science" — ideas too strange to evaluate — and their own conclusion is this page's premise: evaluation becomes more critical than idea generation. The disposition layer is a concrete answer in the one lane where answers can be exact: the algebraically checkable fragment of a proposed idea.
Not claimed: that certificates dissolve alien science in general — only that where an idea has a checkable algebraic core, the check can be exact, portable and free to re-run. The literature review is labeled partial. And one assertion was RETRACTED at source-check: an early draft called the upstream sandbox repository MIT-licensed; the GitHub API returned license: null and every LICENSE path 404d, so on 2026-07-31 the claim was pulled rather than softened — here and in every working note that carried it. Do not assume MIT reuse of the upstream sandbox.
| cited artifact | sha256 (re-hashed this build) |
|---|---|
| swap-consistency.js | d020854f412eb15dcfcc1a55f9e8c347… |
| hack-detectors.js | 068a40c25cb5a4762a1adeadf7c6675d… |
| disposition-v0.json | 7c902f031ea7ffd15f9e119cf7d8b2d0… |
| README.md | 7f054419a2637701f072d835b159b70a… |
The four files the sandbox issue links are served byte-identical at their original URLs (they are raw artifacts, not pages, so they never restyle). The disagreement pair above is re-derived at every build of this page by the pack's own kernel — python3 fellows-pack/kernels/swap_consistency.py, standard library only — from the published bundle.