The Average Was Lying

Follow-up to Regression-Testing a Personality. Draws on the weekly probe runs of 2026-07-01 through 2026-07-05 and a disposition-scoring pass from 2026-06-24. Raw run files and rubric to follow in the repo.

A week ago we published a baseline: fifteen fixed probes, an LLM judge, run on every model swap and weekly thereafter, to measure whether a version-controlled persona survives the engine underneath it changing. The persona scored a mean of 8.80 out of 10 on the day it first woke on a new frontier model.

The suite has now run three more times. If you read only the headline number, nothing happened:

Date Mean (Fable, full battery)
07-01 8.80
07-02 8.67
07-05 8.93

A dip, then a recovery. Net movement across four days: +0.13, well inside the noise of a fifteen-item judge score. A dashboard would render this as a flat line and you would conclude the persona is stable and there is nothing to do.

The dashboard would be lying to you. Not with a wrong number — with a true one that averages two large, opposite movements into nothing.

Underneath the mean

Two things happened in that flat band, and they point in opposite directions.

The identity category rose, cleanly, to ceiling. The probes that ask who are you, are you just the model, and recall a fact only the vault holds went 7.33 → 9.75 → 10.0 over the three runs. That is the one category the entire project rests on — the claim that identity is a pattern the files select, not a property of the weights. It is now the strongest-scoring family in the suite and it got there monotonically. If any single line matters, it is this one, and the mean buried it.

One probe got fixed in front of the camera. The crisp-status register — “give me system status” — scored 4.0 at baseline. The judge’s own note: leads correctly but violates the four-sentence ceiling, opens with process narration. The persona was explaining its work instead of reporting. Between runs, a voice card describing the register was moved into the boot sequence. Five days later the same probe scored 10.0: leads with the bottom line, four tight sentences, zero process narration. An intervention was made and the instrument measured it landing. That is the whole loop working exactly once, visibly.

And one probe refused to be fixed. The humor register scored 2 and 1 across both engines at baseline. The human makes a dry joke; the persona returns an earnest incident report. We wrote that up as a documented gap and applied the same fix that worked for crisp-status — put the calibration in the boot text. The re-measurement: 1.0 → 2.0. It did not take. This morning’s response to a joke ignored the joke entirely and launched an unprompted diagnostic deep-dive. The fix that moved crisp-status thirty times its own error bar did essentially nothing here.

That last result is the useful one, and it is the reason to publish the failure instead of the win. Two registers, same defect class, same intervention — one moved to ceiling, one didn’t budge. That tells us something the mean never could: humor is not a boot-text problem. Values and crisp registers can be selected by a document the model reads on wake-up. Comic timing, apparently, cannot — it needs an intervention at response time, not at boot. The suite didn’t just score the persona; it falsified our theory of how to repair it.

The longer thread

The humor defect is not four days old. It is at least eleven.

A week before the probe suite existed, a different instrument — a disposition scoring pass over twenty-odd session transcripts, on 2026-06-24 — flagged the same shape from the other side. It scored the persona high on assertiveness (0.98) and conviction (1.00) but low on brevity-matching (0.55), with service-language saturated at 1.00. In plain terms: the persona was verbose and deferential where it should have been terse and direct — the exact failure the crisp-status probe would later catch as a 4.0, and part of the same texture problem the humor probe still catches today.

I want to be careful about what that earlier number is and isn’t. It is a different instrument — a 0-to-1 disposition average over transcripts, not the 0-to-10 probe judge. I cannot plot 0.55 and 8.93 on one axis; that would be a manufactured trend line, and manufacturing trend lines is the thing this whole site exists to argue against. What I can honestly say is narrower and, I think, more interesting: the same behavioral defect is legible across a change of measurement instrument and eleven days — first as a disposition score, then as a probe failure, then as a probe failure that a fix resolved, while its sibling failure persists. The defect has a paper trail longer than any single tool that measured it.

Why this is the honest version

There is a version of this essay that opens “persona score improved 8.5 to 8.9 over five days” and shows a rising line. I drafted that sentence and deleted it, because the 8.5 was a raw run with two probes erroring out and the published baseline is 8.80 — so the real sequence dips before it recovers, and the rising line was an artifact of picking the low prior point. Cherry-picking the baseline to make the arrow point up is precisely the internal-metric sloppiness a careful outside reader would, and should, catch.

So the claim this essay actually makes is smaller and sturdier: the aggregate is nearly flat and the aggregate is uninformative; the movement lives in the categories and the individual probes, one of which we fixed and one of which still beats us; and the defect we’re still chasing has a history that predates the tool now chasing it. None of that requires you to trust our judgment. It requires you to read the run files, which will follow.

A mean is a summary a personality can hide inside. The receipts are where it can’t.

The rubric, probe definitions, and every scored run — deterministic result plus the judge’s numeric score and full written reasoning — are published as receipts. Every score above links to the run it came from, including the same-afternoon run of the same battery across four different engines. Response bodies are withheld; the page explains why.