The Warm-Up Must Be Related
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.
In plain English
What we asked. Starting a small model on a mix of easier tasks made it much better at a hard one. But was it the easier versions of the same task that helped, or would any varied warm-up do? A recent paper suggested only related tasks help.
What we found. Only related material helped. Easier versions of the same copying task lifted the result from 58 to 79 percent, on every run. An easy task of a different kind, adding instead of copying, did nothing at all. And a warm-up on examples whose answers had been shuffled left the model worse off than if it had just started later.
Why it matters. In practice: warm up on simpler versions of the thing you want the model to learn, not on whatever easy data is to hand. And avoid noisy labels early; the damage lingers.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU,32training runs of6000steps. The design and the kill test were committed (7703ff2) before any run.
Program v2 Bucket A, item A21. Decisive computation: . Output: analysis/warmup_content.py.analysis/warmup_content.json
The question
arXiv 2505.18369 ("Small Models, Smarter Learning: The Power of Joint Task Training") reports that easy tasks trained alongside a hard one cut the capacity it needs by 2-7x, only when they share computational elements: "task compatibility, not mere diversity". Every warm-up in A15-A20 was an easier version of its target, so the record could not tell those apart. Is it related material, or any varied warm-up?
Design: A15's copy task (lag 8, length 16, rate 0.01, 6000 steps, J8's donor seeds). Four arms, the same compute, differing only in the first 1500 steps:
- fixed -- lag
8throughout (A15's arm). - lag-mix -- lags
2/4/6at random (A15's arm). - sum-mix -- modular sums
(2,3)/(2,4)/(2,5)at random: a different operation on the same kind of random stream and the same vocabulary. - noise -- the lag mix with each batch's scored targets permuted: the same inputs and target frequencies, no learnable relation.
Kill test, fixed before execution: lag-mix minus sum-mix includes zero. Anchor, in code: fixed and lag-mix reproduce A15's runs exactly -- held on all eight seeds.
Results
| Arm | Final lag-8 accuracy | Minus fixed, paired | Seeds above fixed |
|---|---|---|---|
| fixed | 0.575 [0.492, 0.657] | -- | -- |
| lag-mix | 0.790 [0.765, 0.815] | +0.215 [+0.135, +0.296] | 8 of 8 |
| sum-mix | 0.587 [0.511, 0.662] | +0.012 [-0.084, +0.108] | 3 of 8 |
| noise | 0.417 [0.397, 0.437] | -0.157 [-0.244, -0.070] | 1 of 8 |
The kill test does not fire: lag-mix minus sum-mix is +0.203 [+0.122, +0.285], lag-mix higher on 8 of 8 seeds. An easier version of the target helps; an easier version of a different task does not. This agrees with 2505.18369's finding on a different kind of model and task.
The noise warm-up does harm, and the harm outlasts it. It ends 0.157 below fixed. That is more than the 1500 lost steps would explain: fixed at step 4500 (the same number of steps on lag 8) is already at 0.543, and the noise arm, with those 4500 steps after its warm-up, ends at 0.417. Training on targets with no relation to the inputs leaves the model somewhere it takes longer to leave than to have started from scratch.
Sum-mix sits between. It trails fixed through most of the run (0.282 against 0.490 at step 2000) and draws level only at the end; with its 4500 steps on the target it ends at 0.587, a little above fixed's 0.543 at the same number of target steps. At matched total compute, it neither helps nor costs.
Mean lag-8 accuracy | Step 1500 | Step 2000 | Step 3000 | Step 4500 |
|---|---|---|---|---|
| fixed | 0.428 | 0.490 | 0.545 | 0.543 |
| lag-mix | 0.048 | 0.437 | 0.608 | 0.733 |
| sum-mix | 0.032 | 0.282 | 0.439 | 0.516 |
| noise | 0.031 | 0.227 | 0.334 | 0.375 |
What it says
With A20 (a few hundred steps suffice) and A17 (mixing throughout does not help), the recipe is now specific: a brief phase of easier versions of the target task itself, then the target. Varied data alone does not do it, and meaningless data does damage. The mechanism reading this favours is that the easy lags teach a copy operation the hard lag reuses; whether a single easy lag would do as well as a mix is still open (A25).
What stands
- A21: kill test does not fire. Related material beats unrelated material by
+0.203[+0.122, +0.285], on every seed. - Unrelated but learnable material neither helps nor costs at matched compute (
+0.012[-0.084, +0.108]). - A warm-up on meaningless targets costs
-0.157[-0.244, -0.070], more than its lost steps.
Limits
- One unrelated task (modular sum), one width, one task family as target. The sum task shares the stream and the need to look back two steps, so "unrelated" here means a different operation, not a different domain.
- The shuffled-target arm keeps target frequencies; it does not separate "no relation" from "contradictory relation".
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- accuracy
- The fraction of answers a model gets right on questions it was not trained on.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
- vocabulary
- The set of distinct symbols a model can read and produce.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.