Open research: how AI models actually learn
What this is for. Training an AI model is expensive, and most of that cost is spent before anyone knows whether it was necessary. This is our ongoing search for ways to train models more cheaply, more efficiently and with fewer wasted runs, and for measurements that tell you early which approaches are worth paying for. It feeds directly into the models we build for clients: what we find here is what we do, or deliberately stop doing, on real work.
We run it in the open, using models small enough that anyone can reproduce every number on a laptop in minutes, and we publish the results and the data behind them here, including the ones that did not go the way we expected.
How to read the language in these records. These pages are a working laboratory notebook: the design, the numbers, the controls and the limits. Because that is what they are, you will read that an approach did not work, that a result did not hold up under a harder test, or that one method was worse than another. Those are findings, not admissions. Ruling an approach out is how the search narrows, and these are the pages that have taught us the most: a technique that looks promising and then fails a control tells you exactly where the next idea has to be different. Nearly everything we now rely on came out of understanding why something else fell short. If you would rather have the story than the record, the blog carries a plain-language version of most of these, written for a reader who has never seen a training curve.
Start with a question, not a record. Each card below is one question we set out to answer, with where it currently stands and the results behind it. Within a question the records are in the order we ran them, which is the order they are best read in: most build on the ones before, and several exist only because an earlier result made a new question askable.
Why we publish it rather than keeping it. We build custom AI models that you own and run yourself, using our own platform, and a model is only worth what its evaluation is worth. So we publish the whole record, not the flattering half: if a supplier cannot show you the experiments that went against their own idea, you have no way to check the ones that went their way. Holding our own work to that standard in public is the strongest evidence we can offer that we will hold yours to it too.
Ongoing research. These are experimental results from active work, not settled conclusions. The numbers are what we measured and the methods are described so you can judge them, but the programme is still running: later experiments here have already overturned earlier readings more than once, and several pages record exactly that. Expect this section to change as the work moves forward.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.
Top findings so far
If you read only one part of this page, read this: the results we think are most useful to anyone training models, ranked, with why each matters. "Settled" means it survived the controls we planned before running it; "Methods" means a lesson about how to measure that applies well beyond this project. Write-ups of the strongest results for outside readers are drafted here.
A brief start on easier versions of a task reaches the solution in about 40% fewer steps
Against the best ordinary recipe we could tune, including a learning-rate warm-up, on fresh runs: about 42% fewer steps at one model size, 46% at twice the size, and the same for subtraction and for bitwise XOR as for addition. It must be an early phase, it appears in two kinds of recurrent network (GRU and LSTM, each against its best rate) but not in a small transformer. It helps when the hard task is hard for ordinary training and the easier versions share its rule, and the harder the task the more it saves (up to 40%); on an easy version, or on a random lookup table with no rule (easy or hard), it cost time, while the sum with the same symbols and model was halved. The rule has to be shared exactly: break a quarter of it and the gain is gone.
Read the resultOn another task, the same trick was a learning-rate warm-up in disguise
Its large gain vanished once ordinary training was tuned, and a standard warm-up schedule matched it. A cheap test every data-curriculum claim should pass.
Read the resultA trained model's inner machinery gives a new model a large head start
Copying a converged model's recurrent weights cuts a new model's time to learn by about 39% of the run, even with its learning rate tuned; about five new models pay for one donor.
Read the resultMost of our speed-up claims had been measured against an untuned baseline
An audit found it across the archive, and the same trap caught us again before a second audit came back clean. The most common way an efficiency result is wrong.
Read the resultGrowing the model on a fixed task makes effects look like they fade
With the task kept equally hard as the model grew, an effect reported as vanishing stayed flat. Scaling claims need the task scaled too.
Read the resultOne training run predicts where a batch-size sweep bends
A quantity measured in a single run lands within a factor of two of the result of a full sweep, saving the sweep when choosing a batch size.
Read the resultA free early-warning signal was there all along
Held-out loss moves before held-out accuracy in every run, at no cost, where a purpose-built probe cost more than half the training budget.
Read the resultEverything, by question
Everything below is one project, asked as 10 questions. Each has a short answer, and behind it the 263 individual results that got us there. Results that rule an approach out are published exactly like the ones that confirm it, and several of the most useful entries are places where a later check of ours revised an earlier conclusion of ours. That is the process working as intended.
Does reshaping a model's internals make training cheaper?
No. The technique does exactly what it claims and the claim does not translate into an efficiency gain -- ordinary training reaches the same place, and the overhead is 23x the measured benefit.
See the resultsCan you tell in advance that a model is about to improve?
Yes, repeatedly and by several independent routes -- a probe on the hidden state, gradient statistics, held-out loss, and which earlier word the prediction leans on. The warnings are real; how far ahead they fire varies a great deal.
See the resultsIs seeing it coming actually worth anything?
So far, no. Both attempts to act on a warning were matched by acting at a random moment, and the detection cost more than the manoeuvre saved. What the warning is worth is a separate question from what makes training cheaper, which has its own page.
See the resultsWhat actually happens at the moment a model learns?
It builds machinery rather than selecting it, working through the task in a reproducible order and trying a simpler wrong rule on the way. The visible training curve cannot tell you which is happening.
See the resultsWhat decides when a model learns, and can you change it?
Settings dominate, data barely matters, and there is a brief window before the jump in which interrupting the model is unusually costly. Timing can be delayed but not brought forward.
See the resultsWhich parts of a model actually matter?
The parts that move most are not the parts that matter, and no single component is required -- the model routes around every freeze. What a part is worth shows up only when you remove it.
See the resultsHow much model does a task need, and what changes when it has more?
The abrupt jump is what spare capacity buys -- it fades smoothly as the model shrinks, long before the model stops working. Capacity and task difficulty act separately, not as a ratio.
See the resultsIs the task we are studying actually hard?
Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.
See the resultsHow do we know our own results are real?
Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
See the resultsWhat actually makes training cheaper?
One thing has worked: stopping part of the training early saved about 7% with no loss of quality. Everything else tested has been matched by a simpler or cheaper method -- and in two cases the clever method was only winning because it was quietly being given more.
See the resultsWondering if your data is AI-ready?
Start with a free readiness scorecard you run yourself. We never see your data.
Get the free scorecard