The Moment Marks Nothing
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
Part of a bigger question: How do we know our own results are real? – Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
In plain English
What we asked. The page before this one concluded that a training shortcut helps if you take it before the model's sudden improvement and hurts if you take it after. That conclusion is withdrawn, by this experiment, on the same day it was published. The earlier grid could not tell the two candidate explanations apart, and we said so in its own write-up, and then preferred one of them anyway.
What we found. This experiment places the switch at a chosen multiple of each individual run's own moment, on a single task, so that the only thing differing between the two sets of runs is the size of the model. The sign of the result differs by model size at three of the four timings, including one clearly after the moment where the smaller model still saves nine per cent. On the smaller model the benefit fades gradually the later you switch, passing straight through the supposedly special moment without a kink, and only turns negative at around twice it. On the larger model the shortcut costs at every timing we tried. So it is model size, and the moment marks nothing.
Why it matters. The part worth publishing is the mistake. The earlier experiment had a quick test written down in advance, it gave the right answer, and we overrode it on the strength of a pattern that fitted every single point we happened to have. That is exactly what a confounded design produces: when two explanations are tied together, they agree on all the evidence that exists, so finding a pattern that explains everything is weak evidence rather than strong. We had even written down the reason to distrust the test - and never checked that the same reason applied to the explanation we preferred instead.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Negative results are kept. Roughly half of what is published here says an idea did not work, including several of our own. Those pages are not failures, they are the output. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 90 training runs, no GPU, no cost.
Program v2 Bucket Q, item Q11. Decisive computation: . Output: analysis/side_of_transition.py. Reproduce with analysis/side_of_transition.jsonpython analysis/side_of_transition.py; --reuse re-derives every endpoint without retraining.
The question
Q8 reported that the sign of J7's batch cut is decided by which side of the run's own transition the switch lands on: six informative arms, six for six. It also said, in its own verdict, that its grid could not prove this: the switch was fixed at a fraction of the budget while the transition moves with width, tying the two explanations together.
This unties them. The switch is placed at a multiple of each run's own control transition, 0.5x and 0.8x before it, 1.25x and 2.0x after: at two widths on one task, so difficulty and scored tokens per step are held fixed and only width varies.
Kill test, fixed before execution: at a matched multiple the sign is the same at both widths.
The anchors are exact
| Here | Q8 | |
|---|---|---|
w48-lag4 control | 10,944 [9,772, 12,116] | 10,944 [9,772, 12,116] |
w48-lag4 transition | 88 | 88 |
w192-lag4 control | 6,080 [5,799, 6,361] | 6,080 [5,799, 6,361] |
w192-lag4 transition | 39 | 39 |
Result: the kill test does not fire, and Q8's reading is withdrawn
| Multiple of the run's own transition | Width 48 | Width 192 | |
|---|---|---|---|
x0.5 (before) | +25.5% | -46.3% | sign differs |
x0.8 (before) | +19.9% | -42.0% | sign differs |
x1.25 (after) | +9.4% | -33.3% | sign differs |
x2.0 (after) | -1.9% | -53.4% | same sign |
The sign differs by width at three of the four matched multiples, including one clearly after the transition where width 48 still saves +9.4%. Side of the transition does not decide the sign. Width does.
And the transition is not a special point at all. At width 48 the benefit decays smoothly with how late the cut happens, +25.5%, +19.9%, +9.4%, -1.9%, with no discontinuity anywhere near the transition, crossing zero somewhere around twice it. At width 192 the cut costs at every timing tested, between 33% and 53%, with no ordering worth reading.
Q8's six-for-six pattern is reproduced and explained. Its w48-lag11 within-cell pair, +20.6% before, -8.8% at +369 steps, sat at roughly x0.6 and x1.97 of that cell's transition of 381. Here, width 48 gives +19.9% at x0.8 and -1.9% at x2.0. That is the same smooth decay, not a sign flip at the transition. Q8 saw two points on a slope and drew a step.
The part worth publishing is the mistake
Q8's preregistered kill test gave the right answer and I overrode it. Its test was "the cut still saves at width 192, lag 4"; it did not, at -42.4%, which points straight at width. I reported that literal outcome and then argued it did not license its conclusion, because the grid tied width to side-of-the-transition, which was true. But the alternative I preferred was tied by exactly the same confound, and I did not apply the objection to it.
A post-hoc pattern that explains every point in a small confounded grid is not strong evidence. It is the most likely thing such a grid produces, because the confound guarantees the two explanations agree on the points that exist. The discipline that catches this already exists in this programme and I did not use it: when you distrust a preregistered test because of a confound, check whether the confound also protects the explanation you are about to prefer.
The kill test was cheap, fixed in advance, and correct. The reinterpretation was free, fitted after the fact, and wrong.
The fixed-schedule arm, for the fourth time
Each seed's own transition against one population number for every seed, at all eight matched pairs: the oracle's advantage is at most 7.0% and under 1.5% in five of eight, with the largest gap going against the oracle's favour in one case. Per-run timing buys nothing here either, which is K2's finding, Q1's, Q5's and now this one.
What now stands
- Width decides whether cutting the batch helps. At width 48 on lag 4 it saves up to
25.5%; at width 192 on the same task it costs33%to53%, at every timing tested. Difficulty and scored tokens per step are held fixed here, so neither is required. - Cut early rather than late, at the width where it works. The benefit decays monotonically and the transition marks nothing.
- Nothing rescues J7's economics. The manoeuvre works only at the small width, needs no instrument to time it, and the instrument in question costs more than it saves (Q9).
Limits
- Two widths.
48and192bracket a sign change somewhere between them; nothing here locates it, and a width sweep is the obvious follow-up. - One task, one cut ratio, one target. Lag 4 throughout,
4xcut,0.9target, all inherited from J7 and fixed in advance. - The oracle needs a control run per seed and is not a deployable method; it exists to place the switch on a chosen side, and the fixed arm beside it is what a practitioner could use.
- Five seeds, and several intervals are wide,
w48-lag4 x2.0at-1.9%[-...]plainly includes zero, so "crosses zero aroundx2" is an eyeball reading of a monotone sequence, not a fitted crossing point. - The transition here is one of Q2's fourteen definitions, imported rather than re-implemented. Q7's rule applies, and the multiples are far enough apart that no comparison sits inside one evaluation interval.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- confound
- A second explanation you did not control for. If bigger models both learn faster and score higher, then 'fast learners score higher' may be entirely about size and not about speed.
- evaluation interval
- The gap between successive checks of a model during training. It is the smallest difference in timing a measurement can resolve, so two results closer together than one interval cannot be told apart.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- post hoc
- Worked out after the fact, rather than decided in advance. We report such checks separately and never let them decide a result, because it is far too easy to find a pattern once you already know the answer.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
- slope
- How steeply one quantity changes as another does. A slope of one between a warning and the event it predicts means the warning shifts exactly in step with the event.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the experiments that went against us, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.