Research record

A Critical Amount Per Step

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: What actually makes training cheaper? – One thing has worked: stopping part of the training early saved about 7% with no loss of quality. Everything else tested has been matched by a simpler or cheaper method -- and in two cases the clever method was only winning because it was quietly being given more.

In plain English

What we asked. Our rule for how learning time depends on the material per training step broke down when there was very little per step. We checked whether it was still the amount of material that mattered there, or how it was packaged.

What we found. It was still the amount: three very different packagings agreed within 7 percent. The curve just bends. Put together, every setting we have measured fits the standard 'critical batch size' formula used for large language models, within about 2 percent. Below a critical amount per step, adding more saves a lot of training steps; above it, very little.

Why it matters. Practical point: there is a sweet spot for how much to put into each training step, and past it you are paying for material that barely speeds anything up.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 48 training runs. The design and kill test were committed (bcf408b) before any run.

Program v2 Bucket T, item T19. Decisive computation: analysis/content_knee.py. Output: analysis/content_knee.json. Post-hoc curve and fit, changing no verdict: analysis/content_curve.py -> analysis/content_curve.json.

The question

T18 found a power law in scored content per step forecasts learning time within 10% from 704 to 15104 per step, and misses by 46% at 176. T19 asks whether that slow-down is still a function of content per step alone: at 176 and 352 per step, three packagings each (batch 16/8/4 at length 16/27/49; batch 32/16/8 at the same lengths), J8's eight receivers.

Learning time follows one curve in material per step, with a critical point
Learning time follows one curve in material per step, with a critical point. Every clean setting from three experiments, grouped by how many scored symbols the model sees per training step. The dashed curve is the standard critical batch size formula from large-scale training, fitted to these points. The formula fits all nine levels within 2.2% on average. Below about 700 symbols per step, more material per step saves a lot of steps; above it, very little.

Kill test, fixed before execution: at either level, the slowest packaging exceeds the fastest by more than 25%. Anchor: batch 16, length 16 reproduces T13. It holds.

Result: the kill test does not fire

Content per stepBatch x lengthStep to 0.5
17616 x 16386.9 [359.4, 414.4]
1768 x 27399.5 [388.1, 410.9]
1764 x 49412.4 [384.0, 440.7]
35232 x 16233.7 [220.1, 247.3]
35216 x 27242.6 [232.2, 253.0]
3528 x 49247.8 [234.9, 260.7]

Spread 1.066 at 176 and 1.060 at 352: content per step still decides, down to batches of four sequences. T18's power law forecast 211 and 179; the observed 400 and 241 are far slower. The law is right about what matters and wrong about the shape.

The shape: a critical content per step (post-hoc)

Collecting every clean cell on this task (T13, T16, T19; length >= 16) gives nine content levels spanning 86-fold. The local exponent flattens steadily:

Content per step176→352352→704704→14081408→28162816→15104
Local exponent-0.73-0.54-0.38-0.32about -0.14

At small content, doubling it cuts the steps by about 40% (most of the extra content is used); at large content doubling it barely helps. That is the shape McCandlish et al. (2018, arXiv 1812.06162) describe for the critical batch size, and their form fits it closely:

steps = 81.3 x (1 + 716 / content per step), median error 2.2% over all nine levels (largest 6%).

On this task the critical content per step is about 716 scored positions: below it, steps fall steeply as content is added; above it, extra content per step buys little. T8's bending batch curve, T13's interaction and T18's miss are all this one curve seen from different sides.

What stands

  • Kill test does not fire. At 176 and 352 per step, packaging changes learning time by at most 6.6%.
  • The whole curve is one function of content per step, and McCandlish et al.'s critical-batch form fits it within 2.2% with a critical content of about 716.
  • Generates T20: McCandlish et al. predict the critical batch from the gradient noise scale, which can be measured without any sweep. Does the measured noise scale land near 716?

Limits

  • The fit is post-hoc, on cell means, over cells from three records. T20 is the out-of-sample test.
  • One task, one rate (0.006), one model.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

batch size
How many examples the model looks at before updating itself once. Bigger batches give a steadier but more expensive update.
critical batch size
The batch size beyond which making batches bigger barely reduces the number of training steps needed. Below it, bigger batches help a lot; above it, they mostly cost more compute.
exponent
The number in a power law that says how strongly one quantity responds to another. A bigger exponent means a steeper response to the same doubling.
gradient
The direction and amount by which each of a model's internal numbers should change to do slightly better. Training is repeatedly following it.
gradient noise scale
How noisy a single example's learning signal is compared with the average signal over all examples. It suggests how many examples per step are worth using.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
post hoc
Worked out after the fact, rather than decided in advance. We report such checks separately and never let them decide a result, because it is far too easy to find a pattern once you already know the answer.
power law
A relationship where one quantity changes by a fixed percentage whenever another one doubles, rather than by a fixed amount. Most scaling results in AI are stated this way.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.