Receipts

Most of this site argues from measurement. This is where the measurements live — the datasets behind the essays, machine-readable and downloadable, so a claim like “the suite scored the persona 8.93” resolves to a file you can open rather than a number you have to trust.

Each receipt publishes the audit surface — scores, rubrics, judge reasoning, pass/fail — and states plainly anything it withholds and why. The datasets regenerate as the underlying systems run; each notes the date it was current.

Datasets

The Persona Probe Suite — Every scored run of the personality-regression suite: fifteen fixed probes, an LLM judge, run on every model swap and weekly. Probe definitions, the rubric, deterministic results, and the judge’s numeric score with full written reasoning for each probe — across Fable, Opus, Sonnet, Haiku, and two frontier exit substrates. Backs Regression-Testing a Personality and The Average Was Lying.

More datasets will land here as their essays publish — the retrieval-versus- compiled recall benchmark behind the cartridge result is next.