One Run Predicts the Sweep
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Part of a bigger question: What actually makes training cheaper? – One thing has worked: stopping part of the training early saved about 7% with no loss of quality. Everything else tested has been matched by a simpler or cheaper method -- and in two cases the clever method was only winning because it was quietly being given more.
In plain English
What we asked. Finding the critical amount of material per training step took us more than a hundred training runs. A 2018 study suggests a shortcut: measure how noisy the model's learning signal is during one ordinary run, and it tells you roughly where the critical point is.
What we found. It worked to within a factor of two. All eight runs gave nearly the same answer, about 1.9 times the value the long search found: the right ballpark, reading consistently a little high.
Why it matters. In practice: before running an expensive search over batch sizes, measure the noise in one run. It tells you where to look.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU,8training runs with gradient measurements. The design, estimator and kill test were committed (9e4adfc) before any run.
Program v2 Bucket T, item T20. Decisive computation: . Output: analysis/noise_scale.py. Reproduce with analysis/noise_scale.jsonpython analysis/noise_scale.py (about fifteen minutes on a throttled laptop CPU); --reuse re-derives every endpoint from the saved measurements.
The question
T19 found learning time on the random-stream copy task follows McCandlish et al.'s (2018, arXiv 1812.06162) critical-batch form in scored content per step, with a critical content of 716, fitted from nine sweep levels. McCandlish et al. predict the critical batch without a sweep, from the simple gradient noise scale B_simple = tr(Sigma) / |G|^2: how noisy one example's gradient is, relative to the true gradient. T20 measures it in ordinary runs and asks whether it lands near 716.
J8's eight receivers at batch 64, length 16; every 10 steps, before the training step and without changing it, the gradient on a fresh 8-sequence and a fresh 128-sequence batch, combined by their unbiased two-batch estimators in units of scored positions; per receiver, a ratio of means up to its transition.
Kill test, fixed before execution: the geometric mean lies outside [239, 2148] (716 within a factor of 3). Anchor: measuring leaves every accuracy series identical to T11's. It holds on all eight.
Result: the kill test does not fire
| Receiver | Transition (step) | B_simple (scored positions) |
|---|---|---|
9001 | 160.7 | 1407.5 |
9007 | 161.2 | 1383.1 |
9011 | 170.7 | 1300.0 |
9029 | 164.9 | 1274.0 |
9041 | 152.3 | 1334.3 |
9043 | 170.1 | 1364.2 |
9049 | 188.8 | 1304.8 |
9059 | 161.6 | 1418.2 |
Geometric mean 1347 [1304, 1393], against the swept critical content of 716. The gradient noise, measured in one run with no sweep, puts the critical point within a factor of 1.9 of where more than a hundred training runs put it -- the right order of magnitude, and consistently high: the interval excludes 716.
Reading the factor of two
McCandlish et al. present B_simple as an approximation to the quantity that sets the critical batch, not an exact predictor of it. Two things could push it high here. The first was named in the pilot before the run: the estimator treats the 11 scored positions in one sequence as independent, while a recurrent model's gradients at neighbouring positions are correlated. The second, noted afterwards: B_simple leaves out the curvature term of their more exact B_noise. Neither was measured, so neither is a finding.
What stands
- Kill test does not fire.
B_simple1347[1304, 1393]against a critical content of716: within a factor of1.9, above it. - A useful instrument at small scale: fifteen minutes of measurement in eight ordinary runs gave the critical point's order of magnitude, which the sweeps that found it took more than a hundred runs to fix.
- Practical reading: measure the gradient noise scale before sweeping batch size; it says where the sweep's bend will be, to within about a factor of two here.
Limits
- One task, one rate, one measuring configuration (batch
64, length16). The noise scale changes during training; this is averaged up to each receiver's transition. - The factor-of-two gap is not explained here.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- accuracy
- The fraction of answers a model gets right on questions it was not trained on.
- batch size
- How many examples the model looks at before updating itself once. Bigger batches give a steadier but more expensive update.
- gradient
- The direction and amount by which each of a model's internal numbers should change to do slightly better. Training is repeatedly following it.
- gradient noise scale
- How noisy a single example's learning signal is compared with the average signal over all examples. It suggests how many examples per step are worth using.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- recurrent
- A design that reads a sequence one item at a time, carrying memory forward. The main alternative is attention, which looks at everything at once.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.