Research record

The Burst Works Only at the Start

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.

In plain English

What we asked. A hundred steps of easier examples at the start of training made a small model much better at a hard task. Is that because it learned something useful from those examples, or because they steered training early on? If it learned something, the same examples should help whenever they come.

What we found. They helped only at the start. The same hundred steps placed a sixth or half of the way into training did nothing lasting; at the start they lifted the result from 58 to 82 percent, on every run.

Why it matters. Early training has a window when a little of the right data changes where a model ends up. Once it closes, the same data no longer matters. If you are going to shape a model's early data, do it first.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 32 training runs of 6000 steps. The design and the kill test were committed (b18883d) before any run.

Program v2 Bucket A, item A26. Decisive computation: analysis/warmup_timing.py. Output: analysis/warmup_timing.json.

The question

A23 found a 100-step burst of easier copy lags at the very start of training gives the whole warm-up gain (+0.242), and that the gain opens thousands of steps after the burst ends. Two readings: the burst supplies a skill the hard task reuses, which should work wherever it is placed; or it sets the path training takes, in which case its timing should matter. Does the same burst work later?

The same 100 easier steps help only at the very start
The same 100 easier steps help only at the very start. Every model trains for 6,000 steps on the hard task, except for one burst of 100 steps on easier versions of it. The burst comes at the very start, at step 1,000, or at step 3,000. The grey line has no burst at all. At the start, the burst lifts the final result from 58% to 82% on every run. Later, the identical burst knocks accuracy down briefly and leaves no lasting benefit. There is a window early in training when a little of the right data changes where the model ends up.

Design: A15's task, rate 0.01 and seeds (J8's donor seeds), 6000 steps. A 100-step burst of lags 2/4/6 at steps 1-100 (A23's arm), 1001-1100 or 3001-3100, every other step lag 8; plus fixed. The same compute in every arm. Kill test, fixed before execution: the early burst minus the step-1001 burst includes zero. Anchor, in code: fixed and the step-1 burst reproduce A23's runs exactly -- held on all eight seeds.

Results

Burst placed atFinal lag-8 accuracyMinus fixed, pairedSeeds above fixed
none (fixed)0.575 [0.492, 0.657]----
steps 1-1000.817 [0.772, 0.861]+0.242 [+0.163, +0.321]8 of 8
steps 1001-11000.585 [0.506, 0.663]+0.010 [-0.108, +0.129]3 of 8
steps 3001-31000.519 [0.448, 0.590]-0.056 [-0.130, +0.019]2 of 8

The kill test does not fire. The early burst beats the same burst at step 1001 by +0.232 [+0.181, +0.283] and at step 3001 by +0.298 [+0.236, +0.359]. The same 100 steps of the same material help only at the start. Placed at step 1001, the burst knocks lag-8 accuracy from 0.369 to 0.099 for a moment, the model recovers by step 1500, and it ends where fixed does.

When the window is open. The fixed arm is not idle before step 1000: its lag-8 accuracy is already 0.204 at step 100 and first reaches 0.3 between steps 390 and 690 depending on the seed. The window closes somewhere between step 100 and step 1000, while the model is committing to the solution it will refine for the rest of the run.

What it says

The burst works by setting the path, not by supplying a skill. A skill supplied at step 1001 would have been as reusable as one supplied at step 1; it was not reused at all. With A21 (the material must be related), A23 (a hundred steps suffice) and A25 (one nearby easier version does most of it), the picture is of a critical period at the start of training, during which a small amount of easier, related data decides which solution the model goes on to find.

This meets the programme's earlier critical-period records from the other side. Those found that disrupting training in an early window does lasting damage on this delayed-copy task (L7) and that the window survives at width 192 on a matched ladder (P12's second case). A26 finds that helping in the early window does lasting good, and helping later does nothing. A28 locates the close of the window (bursts at steps 101, 301, 601).

What stands

  • A26: kill test does not fire. The early burst beats the step-1001 burst by +0.232 [+0.181, +0.283], on every seed.
  • A burst at step 1001 or 3001 does not help (+0.010, -0.056, intervals including zero).
  • The window closes between steps 100 and 1000 on this task at rate 0.01.

Limits

  • Three placements; the window's close is not located (A28).
  • One task, one width, one rate; the window is likely to scale with the task's learning time (see the standing rule that a schedule carries its task's clock).
QUALIFIED 2026-09-28 by A29: the baseline's learning rate was never tuned. This record ran at rate 0.01. At 0.005 the plain lag-8 run reaches 0.763 at step 6000 (against 0.575 at 0.01), and the 1500-step warm-up adds nothing (-0.034 [-0.104, +0.035]). The within-rate comparisons above stand as measured; any reading of them as an efficiency gain over the best plain alternative is suspended until A31 tunes both arms. Text and numbers above unchanged.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

accuracy
The fraction of answers a model gets right on questions it was not trained on.
baseline
The thing you compare against. A result without one is not a result.
critical period
A stretch of time during which something has to happen for development to proceed normally. Borrowed from biology, where it describes windows in which a young brain must receive certain input.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
learning rate
How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
width
How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.