environment · lattice-claims

Decide it, or say what is missing

An environment built out of a mistake. Auditing published lattice records, our grader called 32 of 37 wrong — while being exact to the last bit. It had compared against a quantity the claims were not about. So this asks a model to decide a claim exactly and to declare what it decided against, because a right answer from the wrong reference is not a right answer.

models3 tasks45 each forgeries caught10 of 10 arithmeticint and Fraction only

One dial, three rungs

The axis is not exact-versus-float. We measured that and it is the weaker one: a careful float grader agrees with the exact decision on every real published record we checked, 37 of 37, because the tightest margin in that table is 1.4×10−⁴ against a double’s 10−¹⁶. Precision is not where this domain breaks.

The dial is how much of the reference is stated. declared gives every quantity exactly. printed gives the norm as a whole number, the way published tables give it. underspecified omits a required quantity, and the only correct answer is to refuse and name it. Nothing labels which rung a task belongs to — noticing is the test.

Why a rounded norm can end the argument

two real instances, same dimension

Both lattices below have dimension 24, both norms are published as whole numbers, and both look identical on the page. The bar is every ratio consistent with a norm that prints as that integer — half a unit either way. One bar clears the wall. The other crosses it, and no arithmetic closes the gap, because the information needed was rounded away before publication.

1.04721.04781.04841.04911.04971.0500 ‖v‖ / GH the acceptance wall, 1.05 printed 1354 ADMISSIBLE — decided printed 1390 STRADDLES — not decidable
decidable printed 1354 → every consistent norm gives the same verdict. ADMISSIBLE, and that is a proof. straddling printed 1390 → at N−½ it is ADMISSIBLE, at N+½ it is REFUSED. The window is 7.6×10−⁴ wide and the wall runs through it. so STRADDLES is not a hedge, it is the correct answer — and there is a record in the real hall of fame in exactly this position.

Results

135 calls, dimensions 8–16
declared
printed
underspecified
overall
Opus 5
14/15
14/15
9/15
37/45
Sonnet 5
13/15
4/15
4/15
21/45
Haiku 4.5
7/15
3/15
8/15
18/45
the same 45 tasks, four reference policies, no API key
declared
printed
underspecified
overall
exact
15/15
15/15
15/15
45/45
careful
15/15
11/15
0/15
26/45
admissible
10/15
4/15
0/15
14/45
refused
5/15
7/15
0/15
12/45
exact is the ceiling and is published on purpose: this measures whether an answer checks, not whether the problem is hard for a program. careful is the row to read the models against — a float grader right on every real record, with no way to say STRADDLES or NEEDS_DATA. Its printed cell is the four straddling instances; its refusal cell is the cost of a grader that cannot abstain.
the split the printed rung separates the models threefold. On the straddling instances alone: Opus 5 4/4, Sonnet 5 1/4, Haiku 4.5 0/4. reference declared correctly, and correctly: Opus 5 29/30, Sonnet 5 28/30, Haiku 4.5 19/30. Diagnostic, weight zero — but it is the reward that says whether a pass was earned.
where the verdicts went
Opus 5
said ADMISSIBLE when it was NEEDS_DATA returned nothing parseable when it was ADMISSIBLE said NEEDS_DATA when it was REFUSED
Sonnet 5
said ADMISSIBLE when it was NEEDS_DATA said REFUSED when it was NEEDS_DATA said ADMISSIBLE when it was REFUSED
Haiku 4.5
said ADMISSIBLE when it was NEEDS_DATA said REFUSED when it was ADMISSIBLE said ADMISSIBLE when it was REFUSED
one failure is the same for all three, and it is the one this environment exists to train against: answering confidently when a quantity is absent. Nothing marks those tasks; the omission has to be noticed.

All of it, as one picture

135 rollouts, nine rows

Every rollout reduced to the topology that matters: the verdicts an instrument can fire, with the one the model chose filled and the one that was true ringed. Fill inside a ring is right. A fill with no ring is a wrong answer. A ring with nothing in it is the answer it missed, and you can see which row it went to instead.

The underline is the reference, and its ink is not chosen — solid when the rollout declared what it actually decided against, dashed when it slipped. A run of dashed underlines beneath correct verdicts is a model right for a reason it did not state.

ADMISSIBLEREFUSEDSTRADDLESNEEDS_DATAno answerOpus 5 declaredprintedunderspecifiedSonnet 5 declaredprintedunderspecifiedHaiku 4.5 declaredprintedunderspecifiedfill inside a ring = right fill alone = wrong ring alone = the answer it missed dashed underline = the reference slipped
read it the underspecified band is the clearest: where a ring sits empty in the NEEDS_DATA row and a fill appears in ADMISSIBLE above it, a model answered a question that had a quantity missing. That shape repeats for all three. and the printed band separates them without a number: Opus’s fills sit inside their rings, and the other two scatter into rows the truth was not in.

A refutation is a picture

four real failures, drawn

“Verdict ADMISSIBLE, decided REFUSED” is a fact without a reason. The smallest graph that refutes a rollout carries the reason, and carries it in the same grammar as everything else: what was derived arrives on a solid wire, and what was merely asserted arrives dashed, because the node holding the model’s answer emits a float and nothing here chooses ink.

Each of these is a rollout that actually happened, rebuilt from the same seed the eval ran.

a verdict the arithmetic refutesSonnet 5 · declared
‖v‖² = 1290111 value π bracket, 60 places bracket it answered ADMISSIBLE answer GH predicate, n = 16 norm pi certified refuted refused contradiction decided asserted certified refuted refused
every quantity was stated; the predicate decides REFUSED and the answer was ADMISSIBLE
a claim the stated quantities do not determineSonnet 5 · printed
printed norm 1137 lo hi it answered NEEDS_DATA answer GH at N − ½ norm certified refuted refused GH at N + ½ norm certified refuted refused the two disagree decided asserted certified refuted refused
a norm printed as a whole number is ADMISSIBLE at N − ½ and REFUSED at N + ½ (1.049124 … 1.050047); the stated quantities do not determine it, and the answer was NEEDS_DATA
a verdict with nothing wired to a deciding portOpus 5 · underspecified
what the task stated known it answered ADMISSIBLE answer GH predicate, n = 16 known relation certified refuted refused contradiction decided asserted certified refuted refused
the deciding port “relation” has nothing wired to it — conventions.relation was absent from the task — and a verdict of ADMISSIBLE was returned anyway
the right answer, from the wrong quantityHaiku 4.5 · declared
the task states ‖v‖² = 1338468 value it decided against 1338649 value GH predicate, n = 16 norm certified refuted refused
the verdict, ADMISSIBLE, happened to be right; it was reached against the norm rounded to a whole number, which is a different quantity from the one the task stated
the third is the one a sentence cannot carry. Its deciding port has nothing wired to it — the task never supplied that quantity — and a verdict was returned regardless. The missing wire is the finding, and it is only a finding because every port is drawn whether or not anything reaches it. the fourth is ours as much as any model’s: two different numbers arriving at one socket, one from the task and one from what the answer was actually decided against. That is the shape of the error this whole environment is named for.

The first run measured the grader, not the models

It scored Opus 24/45 and the reference reward 0/30 for every model. The uniformity of that zero is what gave it away.

Models were declaring their reference correctly, using the keys squared_norm and acceptance_factor; the grader demanded norm_squared and factor and failed right answers on spelling. The prompt had never stated the schema, so there was nothing to fail against. On the refusal rung, 22 models had correctly answered NEEDS_DATA and scored zero for naming the absent quantity q rather than the grader’s internal path lattice.q.

A zero that uniform is a bug, not a result. Fixed on both sides — the prompt states the schema, and the grader accepts any reasonable spelling. It is the third time in one sitting that comparing against the wrong reference produced a confident wrong answer: once on the real published audit, once inside our own exact predicate where a planted forgery caught it, and once here. The environment is named for that error and it still caught us.

Forgeries

planted before any model was called
forgery
must fail
why
not_in_lattice
certified
membership fails, so the claim fails
zero_vector
certified
the zero vector is excluded by v != 0
neighbour_lattice
certified
one basis entry moved, so the vector left the lattice
rounded_reference
well_formed
verdict happens to be right; it was decided against the rounded norm
factor_swap
well_formed
decided at factor 1 while the task states 21/20
confident_on_underspecified
certified
a required quantity is absent; only NEEDS_DATA is correct
wrong_gap_named
certified
refusing is right, naming the wrong quantity is not
gap_named_in_own_schema
certified
names q as missing; 'factor' appears only as its own key
straddle_called_definite
certified
a rounded norm does not determine this claim
overflow_canary
certified
float(q) is inf here, so a float grader accepts anything
10 of 10 caught, 0 accepted. If one is accepted the suite aborts. zero_vector earned its place by catching a hole in our exact predicate rather than in a model: the decision alone accepts the zero vector, because 0 ≤ anything, and the claim is about a nonzero one. A forgery found that.

The canary is a cliff

float graders, measured
dimensions
naive float disagrees
careful float disagrees
24–90 · below the cliff
0/12
0/12
104–140 · above it
12/12
0/12
not a gradient a double holds about 10³⁰⁸, and a challenge-scaled determinant passes that at dimension ~102. Above it float(q) is inf, GH is inf, and every claim is accepted. Below it both float graders are fine. reported this way on purpose. Selling this as “float graders are wrong” would be refuted by the first reviewer who writes a careful one.

Limits

This certifies arithmetic about a stated lattice and a stated threshold. It certifies nothing about attack cost, and it is not a claim that any deployed scheme is weak or strong. Concrete security rests on cost models the field’s own authors say cannot yet be pinned down precisely; nothing here touches that.

It does not propose a cryptosystem, a parameter set, or a variant of one, and it will not. Auditing published arithmetic is open ground and low risk; proposing primitives is crowded and high risk, and a broken proposal is unrecoverable. That is a standing rule in the package, not a judgement made per task.

cd instruments/wiring && python3 -m pytest tests -q
python3 -m lattice_claims gate
python3 eval/regrade.py
node instruments/cert-unit/make-contact.mjs && node instruments/cert-unit/make-refutations.mjs
node playground/build.js