A Learning-Rate Warm-Up Does the Job
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.
In plain English
What we asked. Our easy-data warm-up only helped in the first hundred steps, and only when training took large steps. That is exactly the problem a standard trick called a learning-rate warm-up solves: start with small steps and ramp up.
What we found. The standard trick reached 89 percent, the same as our best data warm-up and better than the easy-task mix. No special data needed.
Why it matters. Before crediting a clever early-training trick, compare it with the ordinary warm-up schedule almost everyone already uses.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU,32training runs of6000steps. The design and the kill test were committed (94874a5) before any run.
Program v2 Bucket A, item A33. Decisive computation: . Output: analysis/warmup_vs_rate_warmup.py. Descriptive post-hoc against A30: analysis/warmup_vs_rate_warmup.json and its output.analysis/warmup_vs_rate_warmup_posthoc.py
The question
At rate 0.01, a 100-step burst of easier copy lags at the very start rescues training (A23); delayed by 100 steps it does nothing (A28); and halving the rate rescues the plain run with no burst at all (A29). One account: the first high-rate steps on the hard task do lasting damage. The standard remedy for that is a learning-rate warm-up -- start the rate small and raise it over the first few hundred steps. Is the data warm-up doing that job?
Design: A15's task, rate 0.01 and seeds (J8's donor seeds), 6000 steps, all on lag 8 except the data arm. Fixed; rate warm-up rising linearly from 0.01/L to 0.01 over the first L = 100 or 375 steps; data burst -- A23's 100 steps of lags 2/4/6. Kill test, fixed before execution: the data burst minus the better rate warm-up includes zero -- the two do the same job. Anchor, in code: the fixed arm through this script's own loop, and the data burst, reproduce A23 exactly -- held on all eight seeds.
Results
Arm (rate 0.01) | Final lag-8 accuracy | Minus fixed, paired |
|---|---|---|
| fixed | 0.575 [0.492, 0.657] | -- |
rate warm-up over 100 steps | 0.750 [0.657, 0.842] | +0.175 [+0.126, +0.224] |
rate warm-up over 375 steps | 0.894 [0.871, 0.916] | +0.319 [+0.240, +0.399] |
data burst, 100 steps of the mix (A23) | 0.817 [0.772, 0.861] | +0.242 [+0.163, +0.321] |
The kill test does not fire -- in the direction that matters least for the warm-up. The data burst minus the 375-step rate warm-up is -0.077 [-0.112, -0.043]: the interval excludes zero because the ordinary learning-rate warm-up does better than the data burst, not worse. The two do not do "the same job" only in the sense that the rate warm-up does more of it.
Against the thread's best data warm-up (descriptive, post-hoc). A30's 100 steps of lag 5 reached 0.892 on the same seeds at the same rate; the 375-step rate warm-up reaches 0.894. Paired, the burst minus the rate warm-up is -0.002 [-0.035, +0.032]: indistinguishable.
What it says
The data-warm-up thread's effect is what a standard learning-rate warm-up does, and a learning-rate warm-up does at least as much of it. On this task, at this rate, the first high-rate steps on the hard task do lasting damage; easier data in those steps softens them (why is not measured here), a smaller rate avoids them (A29), and a rate that starts small and rises avoids them best. None of this needs special data. A31, which finished alongside, agrees from the other side: with both arms tuned, the data warm-up's edge over the best constant rate is +0.028 [-0.025, +0.080], and its best rate is higher than the plain run's -- what a warm-up is for.
That settles the practical question the thread was asking. The plainest alternative that uses the same inputs -- a standing rule here -- is a rate warm-up, which every modern training recipe already has, and it matches or beats every data warm-up tried. The within-rate findings (A21's "must be related", A27's sweet spot, A28's first hundred steps) remain true descriptions of how easier data behaves in that role; they are not a route to efficiency beyond what a rate schedule gives. The outside-venue draft stays on hold; if it is revived, it is as a methods note on this chain of controls.
What stands
- A33: kill test does not fire, reversed. A
375-step learning-rate warm-up beats the100-step data burst by0.077[0.043, 0.112]and matches the best data warm-up (0.894against A30's0.892). - The data warm-up substitutes for a learning-rate warm-up; it does not beat one.
Limits
- Two warm-up lengths, linear only, at one peak rate. A31 (running) tunes the constant rate; whether any data warm-up adds anything on top of a rate warm-up is A35.
- One task. A32 (queued) tunes the modular-sum result.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- accuracy
- The fraction of answers a model gets right on questions it was not trained on.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- learning rate
- How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
- post hoc
- Worked out after the fact, rather than decided in advance. We report such checks separately and never let them decide a result, because it is far too easy to find a pattern once you already know the answer.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.