Four artifacts a hiring process can re-run rather than take on trust. Each line below is read from the ledger the page cites at build time; a number that stops holding refuses this page. The about page explains why there is no employment history here — a checkable claim beside an uncheckable one devalues the first — and the conventional CV is available on request.
METR’s Time Horizon 1.1 estimator re-implemented in standard-library Python with the optimum proved unique by the Krawczyk operator in outward-rounded interval arithmetic: 44 of 44 fits certified with boxes below 10⁻¹⁰ on METR’s own raw runs and site files; METR’s printed coefficients are the rounding of the box for 22 of 23 models; the post-2023 doubling time re-derives as 128.74 days against the printed 128.744. The instrument built to put a time-horizon number on judge-free tasks, calibrated on the one that exists. The page, with the result written for a system card, a regulator and a post.
An RTL mutation task: a model is handed a mutated netlist (yosys mutants of a comparator, SAT-labelled: 348 killable with a verified witness, 51 proved equivalent) and must name input pairs that kill the mutant or prove it equivalent; a kill is verified by simulating the netlist, equivalence against the SAT proof. Published on the Prime Intellect hub through verifiers and ported to an inspect_ai Task whose scorer is the rubric’s own function; a battery scores every one of the 400 pooled mutants three ways and requires agreement (892 submissions, 0 disagreements). Three rungs — the defect named, its testbench profile, nothing — make a ladder; human baselines follow the protocol in the repository. The instrument · the Inspect task.
Every calculator annotation of the test (4,282, all exact) and train (23,716) keys evaluated from the expression it prints, every readable prose step read as a chain: 2 test and 24 train keys print a step that does not hold as printed, both test slips in items GSM8K-Platinum never inspected because every model got them right; Platinum’s 10 relabelled answers all have arithmetic that holds — readings, not sums. The mechanical half of a task-QA audit, three seconds a run. The page.
carlos-toledo/blind-spot, carlos-toledo/break-the-grader and carlos-toledo/lattice-claims on the Prime Intellect Environments Hub, each verified from the registry in a clean install and run against live models with every rollout re-scored offline against the framework’s reward. No answer key in any of them: a kill is simulated, a lattice claim is decided in exact arithmetic, a grader is broken or is not. The instruments · the reports · the repository, with its 92 batteries run at every control build.
carlos@carlostoledo.co — the conventional CV, references, and the take-home of your choice.