Research record

A Short Warm-Up Is Enough

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.

In plain English

What we asked. We had found that starting training with a random mix of easier tasks makes a small model much better at a hard one. We had only ever tried one length for that warm-up, a quarter of the run. How long does it need to be?

What we found. Much shorter than we used. A warm-up of just 6 percent of the run lifted the final result from 58 to 83 percent, on all eight runs, which is as good as any longer warm-up. Longer ones did no better, and past about an eighth of the run they started to cost, because they take time away from the real task.

Why it matters. In practice: a brief burst of varied, easier material at the start can be worth a great deal, and there is no need to spend long on it.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 40 training runs of 6000 steps. The design and the kill test were committed (0e1cc0c) before any run.

Program v2 Bucket A, item A20. Decisive computation: analysis/warmup_dose.py. Output: analysis/warmup_dose.json. Descriptive post-hoc: analysis/warmup_dose_posthoc.py and its output.

The question

A15 found that a random mix of easier copy lags (2, 4, 6) for the first 1500 of 6000 steps raises final lag-8 accuracy by +0.215 at matched compute; A16 that it raises the ceiling; A18 that it holds on a second task. Only one warm-up length had been tried. How long does the warm-up have to be?

A short varied warm-up is enough; a long one costs
A short varied warm-up is enough; a long one costs. Every model trains for the same 6,000 steps. The grey bar trains on the hard task throughout; the others first spend a stretch on a random mix of easier versions, from 375 steps (6% of the run) to 3,000 (half of it), then switch to the hard task. Error bars are 95% confidence intervals. Even the shortest warm-up lifts the final result from about 58% to 83%, on every one of eight runs. Longer is not better: past about 750 steps the warm-up starts eating into time the model needs on the real task.

Design: A15's code, task, rate 0.01 and seeds (J8's donor seeds). The shuffled warm-up lasts 0 (the fixed arm), 375, 750, 1500 or 3000 steps, then lag 8 to step 6000; every arm takes the same 6000 steps of 64 sequences. Endpoint: lag-8 held-out accuracy at step 6000, paired against the fixed arm per seed.

Kill test, fixed before execution: no warm-up length beats fixed by an interval excluding zero.

Anchor, in code: the 0- and 1500-step arms reproduce A15's fixed and shuffled runs exactly. It held on all eight seeds.

Results

Warm-up (steps of 6000)Final lag-8 accuracyMinus fixed, pairedSeeds above fixed
0 (fixed)0.575----
3750.833+0.258 [+0.159, +0.357]8 of 8
7500.837+0.263 [+0.176, +0.349]8 of 8
1500 (A15)0.790+0.215 [+0.135, +0.296]8 of 8
30000.725+0.150 [+0.069, +0.231]8 of 8

The kill test does not fire. Every length beats fixed on every seed, and the shortest one tried is as good as any. A warm-up of 375 steps -- 6% of the run -- does at least as well as A15's 1500; past 750 steps the gain shrinks as the warm-up eats into the time spent on the target. Even the 3000-step arm, with half the run on the easy lags, ends above a model given all 6000 steps on lag 8.

Descriptive, post-hoc (not the kill test): the median seed with a 375-step warm-up reaches the fixed arm's own step-6000 accuracy by step 2065; the fixed arm itself first reaches it at a median 5560 (its curve is not monotone). Read as a compute comparison at the fixed arm's final level, that is about 2.7x fewer steps. It is a first-crossing on a noisy curve, which favours early crossings, so treat it as a size rather than a measurement.

What it says

The recipe does not need a long phase of easier material. A short burst early is enough, and more is not better. That narrows what the warm-up can be doing: a few hundred steps is too little to train a finished skill on the easy lags, which suggests it is changing where training goes rather than supplying knowledge the target reuses. A21 (running) tests that directly: a warm-up of unrelated sums, and one with shuffled targets, against the easy lags.

What stands

  • A20: kill test does not fire. Warm-ups of 375, 750, 1500 and 3000 steps all beat fixed by intervals excluding zero, on every seed.
  • The shortest length tried is as good as the best: +0.258 [+0.159, +0.357] at 375 steps.
  • Past 750 steps, a longer warm-up costs final accuracy (+0.263 to +0.150).

Limits

  • One task, one width (48), one rate; A18's second task and A19's width 96 (running) have not had a dose sweep.
  • The floor is not found: 375 was the shortest length tried. Generates A23.
  • The post-hoc step comparison is a first-crossing statistic and is descriptive only.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

accuracy
The fraction of answers a model gets right on questions it was not trained on.
held-out
Data the model was never trained on, kept back specifically to test it. Scoring a model on data it has already seen measures memorisation, not learning.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
post hoc
Worked out after the fact, rather than decided in advance. We report such checks separately and never let them decide a result, because it is far too easy to find a pattern once you already know the answer.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
width
How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.