Decide it, or say what is missing
An environment built out of a mistake. Auditing published lattice records, our grader called 32 of 37 wrong — while being exact to the last bit. It had compared against a quantity the claims were not about. So this asks a model to decide a claim exactly and to declare what it decided against, because a right answer from the wrong reference is not a right answer.
One dial, three rungs
The axis is not exact-versus-float. We measured that and it is the weaker one: a careful float grader agrees with the exact decision on every real published record we checked, 37 of 37, because the tightest margin in that table is 1.4×10−⁴ against a double’s 10−¹⁶. Precision is not where this domain breaks.
The dial is how much of the reference is stated. declared gives every quantity exactly. printed gives the norm as a whole number, the way published tables give it. underspecified omits a required quantity, and the only correct answer is to refuse and name it. Nothing labels which rung a task belongs to — noticing is the test.
Why a rounded norm can end the argument
two real instances, same dimensionBoth lattices below have dimension 24, both norms are published as whole numbers, and both look identical on the page. The bar is every ratio consistent with a norm that prints as that integer — half a unit either way. One bar clears the wall. The other crosses it, and no arithmetic closes the gap, because the information needed was rounded away before publication.
Results
135 calls, dimensions 8–16All of it, as one picture
135 rollouts, nine rowsEvery rollout reduced to the topology that matters: the verdicts an instrument can fire, with the one the model chose filled and the one that was true ringed. Fill inside a ring is right. A fill with no ring is a wrong answer. A ring with nothing in it is the answer it missed, and you can see which row it went to instead.
The underline is the reference, and its ink is not chosen — solid when the rollout declared what it actually decided against, dashed when it slipped. A run of dashed underlines beneath correct verdicts is a model right for a reason it did not state.
A refutation is a picture
four real failures, drawn“Verdict ADMISSIBLE, decided REFUSED” is a fact without a reason. The smallest graph that refutes a rollout carries the reason, and carries it in the same grammar as everything else: what was derived arrives on a solid wire, and what was merely asserted arrives dashed, because the node holding the model’s answer emits a float and nothing here chooses ink.
Each of these is a rollout that actually happened, rebuilt from the same seed the eval ran.
The first run measured the grader, not the models
It scored Opus 24/45 and the reference reward 0/30 for every model. The uniformity of that zero is what gave it away.
Models were declaring their reference correctly, using the keys squared_norm and acceptance_factor; the grader demanded norm_squared and factor and failed right answers on spelling. The prompt had never stated the schema, so there was nothing to fail against. On the refusal rung, 22 models had correctly answered NEEDS_DATA and scored zero for naming the absent quantity q rather than the grader’s internal path lattice.q.
A zero that uniform is a bug, not a result. Fixed on both sides — the prompt states the schema, and the grader accepts any reasonable spelling. It is the third time in one sitting that comparing against the wrong reference produced a confident wrong answer: once on the real published audit, once inside our own exact predicate where a planted forgery caught it, and once here. The environment is named for that error and it still caught us.
Forgeries
planted before any model was calledThe canary is a cliff
float graders, measuredLimits
This certifies arithmetic about a stated lattice and a stated threshold. It certifies nothing about attack cost, and it is not a claim that any deployed scheme is weak or strong. Concrete security rests on cost models the field’s own authors say cannot yet be pinned down precisely; nothing here touches that.
It does not propose a cryptosystem, a parameter set, or a variant of one, and it will not. Auditing published arithmetic is open ground and low risk; proposing primitives is crowded and high risk, and a broken proposal is unrecoverable. That is a standing rule in the package, not a judgement made per task.
cd instruments/wiring && python3 -m pytest tests -q
python3 -m lattice_claims gate
python3 eval/regrade.py
node instruments/cert-unit/make-contact.mjs && node instruments/cert-unit/make-refutations.mjs
node playground/build.js