The Warning Gets Longer, Not Shorter
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
Part of a bigger question: How do we know our own results are real? – Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
In plain English
What we asked. We have one instrument that genuinely works: a small probe that can tell, some way in advance, that a model is about to start performing. An earlier experiment of ours suggested how much advance notice it gives is unstable across model sizes, which would make it hard to rely on. We had reason to think that earlier experiment was distorted, for the same reason two others had turned out to be, so we redid it fairly.
What we found. It was distorted, and not in the way we predicted. On a fair comparison the warning does not shrink as models get bigger. It gets longer, from about fifteen steps of notice in the smallest model to about forty in the largest. What we had expected was that the measurement would become steady across sizes once the comparison was fair. It did not, because a longer warning inside a similar-length run is still a changing proportion. So our hypothesis failed and the instrument came out looking better than either version of the experiment suggested.
Why it matters. The reason we are drawing attention to a failed prediction is that the two previous rechecks had both come out the way we expected, and two for two is exactly when it becomes tempting to stop checking and start assuming. This one did not follow the pattern. There is also a smaller lesson about how we phrased the test. Our pass or fail condition compared two numbers against each other, and it came out in our favour, but for the wrong reason: the number we were comparing against moved, rather than the number we cared about improving. We have reported that as a failure rather than a pass, because a test that can be flipped by something other than the thing it is testing has not really been passed.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Negative results are kept. Roughly half of what is published here says an idea did not work, including several of our own. Those pages are not failures, they are the output. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 4 widths x 5 seeds plus an unmatched comparison, no GPU, no cost.
Program v2 Bucket P, item P12, third of four cases. Decisive computation: . Output: analysis/lead_matched.py. Reproduce with analysis/lead_matched.jsonpython analysis/lead_matched.py; --reuse re-derives every endpoint.
The question, and the prediction
J5 reported that I2's reframing does not generalise: the decode probe's lead expressed as a fraction of the run varies with width by 76.6%, more than the step count varies (29.3%). The probe is F2's instrument and this programme's only positive one.
J5's own numbers looked like the confound P11 found and P12's second case confirmed: its transition collapses from 323.5 steps at width 16 to 33.9 at width 192, while the lead in steps wanders without trend. So the fraction's instability might be nothing but a shrinking denominator.
The prediction, stated before running: on a ladder where the transition is held roughly constant, the fraction should become stable and I2's reframing is restored.
It was wrong.
Kill test as written: on the matched ladder the lead fraction still varies across widths by more than the lead in steps does. Anchor: the unmatched arm must reproduce J5's width-192 fraction of 0.837, read from its committed output. It does, at 0.825.
What we found
| Width | Lag | Transition | Lead (steps) | Lead (fraction) |
|---|---|---|---|---|
| 32 | 4 | 108.8 | 14.8 | 0.137 |
| 48 | 6 | 119.4 | 14.9 | 0.125 |
| 96 | 8 | 130.7 | 18.5 | 0.142 |
| 192 | 11 | 156.6 | 39.9 | 0.255 |
Spread across widths, as range over mean, which is J5's own measure:
| Quantity | J5, unmatched | This ladder, matched |
|---|---|---|
| transition | ~2.5 | 0.371 |
| lead in steps | 0.293 | 1.141 |
| lead as a fraction | 0.766 | 0.796 |
The kill test does not fire, and reporting that as a positive would be wrong
The kill test asks whether the fraction varies more than the steps do. On the matched ladder it does not, 0.796 against 1.141. By the letter of the test, it does not fire.
That is arithmetic, not evidence. The fraction's spread barely moved: 0.766 unmatched against 0.796 matched. What changed is that the steps spread nearly quadrupled, from 0.293 to 1.141. The comparison flipped because its denominator moved, not because the quantity under test improved.
This is the trap CLAUDE.md already names in another form, a rule whose grid contains its own components as degenerate points, so a tie can be arithmetic rather than evidence. Here the endpoint is a comparison between two spreads, and either can flip it. The honest verdict is that the prediction failed: matching difficulty did not stabilise the fraction.
What actually happens is more interesting
On a matched ladder the probe's lead in steps grows with width. From 14.8 at width 32 to 39.9 at width 192, roughly 2.7x, while the transition rises only 1.44x.
That is the reverse of the shrinking-denominator story. Had the fraction been a pure artefact of run length, width 192's fraction should have fallen to about 0.095 once its run was lengthened. Instead it is 0.255. The fraction is high at width 192 because the lead itself grew, not because the run was short.
So J5's finding survives with its cause changed. The fraction is unstable across widths, and on a difficulty-matched ladder the instability is driven by a genuinely growing lead, not by a collapsing denominator. That is a better position for the instrument than J5 left it in, a lead that grows with model size is useful, if it can be relied on, and it is not a restoration of I2's reframing, which was a claim about stability.
What this does to the pattern
Two of three re-tested cases were the confound. This one is not. That matters more than the result itself: the two-for-two run made it tempting to assume the rest would follow, and this record is why the remaining case is still listed as untested rather than presumed.
J5's record needs no correction banner. Its verdict: that the fraction does not generalise across widths, stands. What this adds is that the mechanism is not the one its own numbers suggested.
Limits
- The ladder is imperfectly matched. Transitions run
108.8to156.6, a spread of0.371, against a target of roughly constant. Width 192 at lag 11 overshot. A tighter rung (lag 9 or 10) would be a fairer test, and is the obvious follow-up. - The conclusion is robust to that mismatch by arithmetic: width 192's transition is
1.44xwidth 32's while its fraction is1.86x, so lengthening its run further cannot bring the fraction into line with the others without the lead also shrinking. - Width 16 is excluded from the matched ladder. At lag 4 it already transitions at
323.5, so matching it would need a task easier than lag 4, changing the task's character rather than its difficulty. - Five seeds per rung, one task family, one lever for difficulty.
- The first execution stored the transition under the wrong key and recorded
None. It is exactly recoverable, since J5 defines the fraction as lead over transition, so no run was repeated, but the raw rows in the committed JSON carry the reconstructed value, not a measured one.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- confound
- A second explanation you did not control for. If bigger models both learn faster and score higher, then 'fast learners score higher' may be entirely about size and not about speed.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- probe
- A small separate model trained to read information out of a bigger model's internals, used as a measuring instrument rather than as a product.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the experiments that went against us, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.