Research record

The Window Is the First Hundred Steps

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.

In plain English

What we asked. A burst of easier examples at the start of training helped a small model, but only at the start. How long does that window stay open?

What we found. Barely at all. Placed at the very beginning, the burst lifted the final result from 58 to 82 percent. Moved just 100 steps later, it did nothing reliable. The model has not settled on its solution by then; something in the very first steps goes wrong without the burst.

Why it matters. Together with our previous page, where a slower learning rate made the burst unnecessary, the likely story is that the first large training steps on the hard task do lasting damage. The next experiments test that directly.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 40 training runs of 6000 steps. The design and the kill test were committed (a500949) before any run. Read with A29, which landed while this ran: the rate used here (0.01) was never tuned, and at 0.005 the plain run needs no rescue.

Program v2 Bucket A, item A28. Decisive computation: analysis/warmup_window.py. Output: analysis/warmup_window.json.

The question

A26 found a 100-step burst of easier copy lags helps at steps 1-100 (+0.242) and not at 1001 or 3001. When does the window close? If it closes over the few hundred steps in which the plain run is committing to its solution, bursts at step 101 or 301 should still help. If only the very first steps matter, they should not.

Only the very first hundred steps matter
Only the very first hundred steps matter. Every model trains for 6,000 steps on the hard task except for one burst of 100 steps on easier versions of it, placed at the very start or a little later. Same compute throughout, at the learning rate the earlier warm-up experiments used. Error bars are 95% confidence intervals. At the very start the burst lifts the result from 58% to 82%; delayed by just 100 steps it does nothing reliable. Whatever goes wrong without it happens in the first hundred steps, and a slower learning rate avoids it too.

Design: A26's code, task, rate 0.01 and seeds (J8's donor seeds), 6000 steps. The 100-step burst at steps 1 (A26's arm), 101, 301 or 601; plus fixed. Kill test, fixed before execution: the burst at step 101 minus fixed includes zero. Anchor, in code: fixed and the step-1 burst reproduce A26 exactly -- held on all eight seeds.

Results

Burst placed atFinal lag-8 accuracyMinus fixed, paired
none (fixed)0.575--
steps 1-1000.817+0.242 [+0.163, +0.321]
steps 101-2000.613+0.038 [-0.088, +0.165]
steps 301-4000.623+0.049 [-0.059, +0.156]
steps 601-7000.600+0.025 [-0.103, +0.154]

The kill test fires. Delay the burst by a hundred steps and the gain is gone. The window is the first hundred steps of training, not the period in which the plain run commits to its solution -- which comes later: the fixed arm first reaches 0.3 between steps 390 and 690.

What it says, with A29

Whatever the burst does, it does it by replacing the first steps of training on the hard task. Those first steps, at rate 0.01, are what send the plain run somewhere worse; A29 found that at rate 0.005 the plain run does not go there (0.763). The simplest account of both: the first high-rate steps on the hard task do lasting damage, and anything that stops them -- easier data, or a smaller rate -- avoids it. That is the job a learning-rate warm-up does, and A33 tests it directly.

What stands

  • A28: kill test fires. A burst at step 101 gives +0.038 [-0.088, +0.165]; the window closes within the first hundred steps.
  • The effect belongs to the first steps of training at rate 0.01. Its value against a tuned plain run is A31.

Limits

  • One rate, one that A29 shows is poor for the plain run.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

accuracy
The fraction of answers a model gets right on questions it was not trained on.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
learning rate
How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.