The Effect Never Faded
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
Part of a bigger question: How do we know our own results are real? – Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
In plain English
What we asked. Six times we had found something useful happening during training, tested it at larger model sizes, and watched it disappear. We had just written that down as a rule to live by. Then we noticed that every one of those tests made the model bigger while leaving the task exactly as it was, which means the model had progressively less to do. This is the experiment that separates the two.
What we found. The effect never faded. We reran one of those tests with the task made harder as the model grew, so the model stayed under similar strain. The benefit is flat across every size: about forty per cent of the model's learning time saved at the smallest size and about the same at the largest. In the original test it had dropped to six per cent. At the biggest size the two versions differ by a factor of fourteen, and the only difference between them is how hard the task was.
Why it matters. This overturns two of our own records published the same day, and rewrites a rule we had written that morning. The rule was worrying about the right thing and measuring it the wrong way: a test that grows the model while holding the task still does not measure what happens at larger scale, it measures what happens when a model has too little to do. Both are worth knowing and only one is what we claimed. The part worth carrying away is why the mistake was convincing. All six results agreed with each other, so the pattern looked solid, and it looked solid precisely because every one of them was produced the same wrong way. A systematic mistake does not look like noise; it looks like a law. What caught it was reading somebody else's paper, which cost no computing time at all. Five of the six still need rechecking and we are not claiming they will all turn out the same way.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Negative results are kept. Roughly half of what is published here says an idea did not work, including several of our own. Those pages are not failures, they are the output. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 2 ladders x 3 widths x 4 seeds, no GPU, no cost.
Program v2 Bucket P, item P11, generated by P4. Decisive computation: . Output: analysis/matched_difficulty.py. Reproduce with analysis/matched_difficulty.jsonpython analysis/matched_difficulty.py; --reuse re-derives every endpoint.
The question
Six results in this programme fade as models grow, and this session added a standing rule off the back of them. Every one holds the task fixed while width grows. A model that outgrows its task reaches its transition sooner, on the fixed task the control transitions at 135 steps at width 24 and at 59 at width 96, so a fixed analysis window covers a far larger share of the run at the largest width, and none of the six separates "the effect faded" from "the instrument ran out of room".
Kill test, fixed before execution: the head start's share of the receiver's own run still falls across a difficulty-matched ladder, and its interval includes zero at the largest width.
Anchor enforced in code: on the fixed-task ladder the head start must reproduce P9. It does.
The two ladders
Difficulty is raised with width by lengthening the task's lag, calibrated so the control run's transition time stays roughly constant.
Fixed task, every previous sweep in this programme:
| Width | Lag | Control transition | Head start | As a share of the run |
|---|---|---|---|---|
| 24 | 4 | 135.0 | -57.5 [-65.5, -49.5] | -42.6% |
| 48 | 4 | 88.8 | -27.5 [-32.1, -22.9] | -31.0% |
| 96 | 4 | 58.8 | -3.8 [-7.7, +0.2] | -6.2% |
Difficulty matched to capacity:
| Width | Lag | Control transition | Head start | As a share of the run |
|---|---|---|---|---|
| 24 | 4 | 135.0 | -57.5 [-65.5, -49.5] | -42.6% [-48.5%, -36.7%] |
| 48 | 7 | 133.8 | -53.8 [-61.4, -46.1] | -40.2% [-45.3%, -35.0%] |
| 96 | 8 | 122.5 | -52.5 [-60.5, -44.5] | -42.8% [-47.0%, -38.6%] |
The kill test does not fire, and not narrowly. On the matched ladder the head start is flat: -42.6%, -40.2%, -42.8%, every interval overlapping every other. At width 96 it is -52.5 steps against the fixed ladder's -3.8: a fourteen-fold difference at the same model size, from nothing but making the task harder.
What this overturns
P9 is wrong as a statement about scale. It reported that the donor head start "closes as models grow" and called it the fifth instance of a pattern. It does not close as models grow. It closes as the model outgrows the task, and those are different claims. A correction banner is appended to P9's record; its numbers reproduce exactly and only their interpretation is withdrawn.
P8's break-even number is not width-limited. P8 found a donor pays for itself across 4.8 receivers at width 48, and P9's banner said that held "at width 48 and nowhere above it". At width 96 with difficulty matched the saving is -52.5 steps, larger in absolute terms than at width 48. The scope withdrawal was itself too strong, and P8's banner is corrected.
And the standing rule written this session needs rewriting, not qualifying. "Width-sweep before building on a result" was the right instinct and the wrong instrument. A sweep at fixed task difficulty manufactures the fade it then reports. The rule becomes: sweep width with difficulty tracking capacity, and treat a fixed-task sweep as evidence about task saturation, not about scale.
What it does not establish
Only one of the six was re-tested. The critical-period window, F2's decode-probe lead and the two K1 gradient alarms were not. They share the confound by construction, all four hold the task fixed, but "shares a confound" is not "is explained by it", and each needs its own matched ladder. O7 is the cheapest and is filed as P12.
This also does not say the head start survives at any scale. Width 96 is still a toy. What it says is that within the range this programme can test, the fade was an artefact of the sweep design and not a property of size, which is a much better position to be in than the one held this morning.
The uncomfortable part
Three of this session's records: P9, O7, and the standing rule, were written with confidence about a pattern that a forty-minute experiment has now shown to be, in at least one case, entirely an artefact of how we swept. The pattern looked strong precisely because every instance was produced the same wrong way, which is what a systematic confound does: it does not look like noise, it looks like a law.
The thing that caught it was reading a paper, not running an experiment. P4 cost no compute at all.
Limits
- One of six results re-tested, on one task family, at three widths, four seeds each.
- The matched ladder holds the control run's transition time roughly constant (
135,134,122), which is an operational stand-in for "difficulty relative to capacity" and not a first-principles definition. A10%drift remains across the ladder. - Difficulty is raised through lag only. Vocabulary and sequence length are untouched, and a different lever might not behave the same way.
- The donor is matured to
1.4xits own transition on each rung, so donor cost rises with difficulty too. P8's economics are therefore not directly re-derivable from this table, the break-even count at width 96 matched needs its own calculation. - Width 96 remains small. Nothing here speaks to frontier scale, and the paper that prompted it is explicit that its own largest model is 85M.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- confound
- A second explanation you did not control for. If bigger models both learn faster and score higher, then 'fast learners score higher' may be entirely about size and not about speed.
- critical period
- A stretch of time during which something has to happen for development to proceed normally. Borrowed from biology, where it describes windows in which a young brain must receive certain input.
- gradient
- The direction and amount by which each of a model's internal numbers should change to do slightly better. Training is repeatedly following it.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- probe
- A small separate model trained to read information out of a bigger model's internals, used as a measuring instrument rather than as a product.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
- vocabulary
- The set of distinct symbols a model can read and produce.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the experiments that went against us, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.