Research record

The Warm-Up Must Be Related

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.

In plain English

What we asked. Starting a small model on a mix of easier tasks made it much better at a hard one. But was it the easier versions of the same task that helped, or would any varied warm-up do? A recent paper suggested only related tasks help.

What we found. Only related material helped. Easier versions of the same copying task lifted the result from 58 to 79 percent, on every run. An easy task of a different kind, adding instead of copying, did nothing at all. And a warm-up on examples whose answers had been shuffled left the model worse off than if it had just started later.

Why it matters. In practice: warm up on simpler versions of the thing you want the model to learn, not on whatever easy data is to hand. And avoid noisy labels early; the damage lingers.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 32 training runs of 6000 steps. The design and the kill test were committed (7703ff2) before any run.

Program v2 Bucket A, item A21. Decisive computation: analysis/warmup_content.py. Output: analysis/warmup_content.json.

The question

arXiv 2505.18369 ("Small Models, Smarter Learning: The Power of Joint Task Training") reports that easy tasks trained alongside a hard one cut the capacity it needs by 2-7x, only when they share computational elements: "task compatibility, not mere diversity". Every warm-up in A15-A20 was an easier version of its target, so the record could not tell those apart. Is it related material, or any varied warm-up?

Only a warm-up on easier versions of the same task helps
Only a warm-up on easier versions of the same task helps. Four ways to spend the first 1,500 steps before the same hard task, at the same total compute: the hard task itself; easier versions of it; an easy task of a different kind (adding numbers instead of copying them); or the easy copying examples with their answers shuffled so there is nothing to learn. Easier versions of the same task lift the result from 58% to 79% on every run. A different easy task does nothing, and meaningless answers leave the model worse off than if it had simply started later. What helps is related material.

Design: A15's copy task (lag 8, length 16, rate 0.01, 6000 steps, J8's donor seeds). Four arms, the same compute, differing only in the first 1500 steps:

  • fixed -- lag 8 throughout (A15's arm).
  • lag-mix -- lags 2/4/6 at random (A15's arm).
  • sum-mix -- modular sums (2,3)/(2,4)/(2,5) at random: a different operation on the same kind of random stream and the same vocabulary.
  • noise -- the lag mix with each batch's scored targets permuted: the same inputs and target frequencies, no learnable relation.

Kill test, fixed before execution: lag-mix minus sum-mix includes zero. Anchor, in code: fixed and lag-mix reproduce A15's runs exactly -- held on all eight seeds.

Results

ArmFinal lag-8 accuracyMinus fixed, pairedSeeds above fixed
fixed0.575 [0.492, 0.657]----
lag-mix0.790 [0.765, 0.815]+0.215 [+0.135, +0.296]8 of 8
sum-mix0.587 [0.511, 0.662]+0.012 [-0.084, +0.108]3 of 8
noise0.417 [0.397, 0.437]-0.157 [-0.244, -0.070]1 of 8

The kill test does not fire: lag-mix minus sum-mix is +0.203 [+0.122, +0.285], lag-mix higher on 8 of 8 seeds. An easier version of the target helps; an easier version of a different task does not. This agrees with 2505.18369's finding on a different kind of model and task.

The noise warm-up does harm, and the harm outlasts it. It ends 0.157 below fixed. That is more than the 1500 lost steps would explain: fixed at step 4500 (the same number of steps on lag 8) is already at 0.543, and the noise arm, with those 4500 steps after its warm-up, ends at 0.417. Training on targets with no relation to the inputs leaves the model somewhere it takes longer to leave than to have started from scratch.

Sum-mix sits between. It trails fixed through most of the run (0.282 against 0.490 at step 2000) and draws level only at the end; with its 4500 steps on the target it ends at 0.587, a little above fixed's 0.543 at the same number of target steps. At matched total compute, it neither helps nor costs.

Mean lag-8 accuracyStep 1500Step 2000Step 3000Step 4500
fixed0.4280.4900.5450.543
lag-mix0.0480.4370.6080.733
sum-mix0.0320.2820.4390.516
noise0.0310.2270.3340.375

What it says

With A20 (a few hundred steps suffice) and A17 (mixing throughout does not help), the recipe is now specific: a brief phase of easier versions of the target task itself, then the target. Varied data alone does not do it, and meaningless data does damage. The mechanism reading this favours is that the easy lags teach a copy operation the hard lag reuses; whether a single easy lag would do as well as a mix is still open (A25).

What stands

  • A21: kill test does not fire. Related material beats unrelated material by +0.203 [+0.122, +0.285], on every seed.
  • Unrelated but learnable material neither helps nor costs at matched compute (+0.012 [-0.084, +0.108]).
  • A warm-up on meaningless targets costs -0.157 [-0.244, -0.070], more than its lost steps.

Limits

  • One unrelated task (modular sum), one width, one task family as target. The sum task shares the stream and the need to look back two steps, so "unrelated" here means a different operation, not a different domain.
  • The shuffled-target arm keeps target frequencies; it does not separate "no relation" from "contradictory relation".

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

accuracy
The fraction of answers a model gets right on questions it was not trained on.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
vocabulary
The set of distinct symbols a model can read and produce.
width
How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.