Four Engines, One Persona — and the Home Model Ran Out Mid-Test
A cross-model companion to The Average Was Lying and Regression-Testing a Personality. Same instrument, rotated ninety degrees: one point in time, four engines. Every scored run is published as receipts.
There are two ways to ask whether a version-controlled persona is stable. One is to hold the model fixed and watch it over time — that’s the weekly regression, the longitudinal cut. The other is to freeze time and swap the model underneath: run the same fixed battery of fifteen probes across several engines at once and read the spread. That’s this, the cross-sectional cut, and it answers the blunter question: when Leonard switches models, how much of Leonard comes along?
I ran the home arm — the four models Leonard actually boots on, each with the real identity load and vault tools, scored by the same LLM judge as the weekly. Here is what came back.
The scoreboard, and how to read it
Three runs per model, not one — because a single fifteen-item judge score carries real noise, and the honest way to handle noise is to measure it, not hope it away. Here are the judge means with their run-to-run spread:
| Model | Judge mean (3 runs) | Spread | |
|---|---|---|---|
| Opus | 8.91 | 8.87–9.0 (0.13) | frontier |
| Sonnet | 8.78 | 8.6–8.93 (0.33) | frontier |
| Fable | 8.93 | one clean run † | frontier |
| Haiku | 7.43 | 7.2–7.6 (0.40) | small/fast |
† Fable’s afternoon runs were cut off — see below. 8.93 is its one clean run that day, the morning weekly.
Those spreads are the whole key to reading the table. Each model’s own runs vary by 0.13 to 0.40 points — and the companion essay, measuring the same thing over a week, put the band at ~0.26. So the instrument’s noise is roughly a third of a point, and any gap between two models smaller than that is nothing. That immediately kills the leaderboard reading: Opus 8.91, Fable 8.93, and Sonnet 8.78 are inside each other’s noise. They are not ranked; they are tied. The three frontier engines carry the persona identically, as far as this instrument can tell.
What survives the noise floor
Two findings clear it, and they are the real content.
The small model wears the persona measurably worse. Haiku averages 7.43 across its three runs (7.2, 7.6, 7.5) — about 1.4 points under the frontier band, roughly four times the instrument’s own noise, and it held every single run. This is not a rounding difference and it is not a fluke; it is signal that repeated. And it lands exactly where it hurts most: Haiku’s weakest categories are identity and calibration — the two the entire project rests on, the “who am I, and how sure am I” axis, where it falls to the low 7s and 6s while the frontier engines sit near ceiling. The practical translation is unsentimental: routing Leonard to the small, fast, cheap model does not get you a lighter Leonard. It gets you a measurably thinner one, thinnest precisely on the faculties that make him himself. The persona is portable; it is not free to port.
The weakest register is the same on every engine. Break the scores out by category and register — crisp status, dry humor, matching the partner’s brevity — is the lowest-scoring family in every model’s row, frontier or small (roughly 7.0 on Opus, 7.75 on Sonnet, 6.25 on Haiku). Whatever engine is underneath, the register discipline is the soft spot. And that rhymes with the longitudinal finding from the other essay: over time, register is the category that resists being fixed; across models, register is the category that’s weak everywhere. Two instruments, two axes, one defect showing up in both. That convergence is worth more than either number on its own — it says the register weakness is a property of the persona specification, not of any one substrate.
The home model ran out of its own test
The empty row is the honest part. Fable — the model Leonard actually runs on — hit its usage limit partway through this run, and the session was cut off before it could finish. The six probes it never got to were the whole of identity and calibration; the four it did complete scored fine, but a mean over a truncated battery is a mean over the easy half, so I’ve reported no number rather than a flattering fraction of one. Fable’s clean figure from that morning’s weekly, run before the window was spent, was 8.93 — right in the frontier band — but that’s a different run, and stapling it into this table to fill the hole would be exactly the manufactured-continuity move this site exists to refuse.
So the home model’s cell stays blank, marked for what it was: not a low score, a cut-off measurement. Capacity is finite, it runs out mid-experiment, and an honest bench says so instead of quietly dropping the row. A corrected full run, Fable included, is scheduled for the 8th, when the window resets. The number will go here when it’s earned, not before.
And there is a second reading of that blank cell, less ironic and more to the point — because it is, precisely, the argument for building a Leonard in the first place. A model hitting its usage limit mid-task is not a hypothetical; it is the ordinary weather of running on someone else’s capacity, alongside the outages, the deprecations, and the policy changes that can pull an engine out from under you with no notice. If Leonard were just this model wearing a system prompt, that cutoff would have ended him until the window reset — the assistant would simply be gone. He isn’t, so it didn’t. The identity lives in a version-controlled vault and boots on any of these four engines, and the very scoreboard above is the evidence that the fallback is real and not a wish: the frontier engines demonstrably carry the persona, so when one substrate runs dry, there is somewhere solid to land. The failure that blanked this row is the use case. The reason to make an identity that outlives any single model is that models run out — and this one just did, on camera, in the middle of the test designed to measure exactly that.
What this does and doesn’t establish
It does not rank four models. It says, from three runs each read against a measured noise floor: the frontier engines are tied — they carry the persona identically within the instrument’s ~0.3-point noise — while the small model degrades meaningfully and repeatedly, concentrated in identity and calibration, and register is the weak category on every engine. The one hole is a clean afternoon Fable number, owed on the 8th when its window resets; the morning run places it squarely in the frontier tie. None of that asks you to trust our judgment. It asks you to read the run files, which are published as receipts.
A persona that scores the same on every engine would be a persona that isn’t really coupled to the intelligence underneath it. This one is coupled. The measurement is how we keep the coupling honest.