Open research: how AI models actually learn
We run a public research programme on how AI models actually learn, using models small enough that anyone can reproduce every number on a laptop in minutes. We publish the results and the data behind them here, including the experiments that went against what we expected.
These pages are the records: the design, the numbers, the controls and the limits. Most of them are about a measurement that misled us before we caught it, which is the same class of failure that makes vendor benchmarks unreliable. If you would rather have the story than the record, the blog carries a plain-language version of most of these, written for a reader who has never seen a training curve.
Ongoing research. These are experimental results from active work, not settled conclusions. The numbers are what we measured and the methods are described so you can judge them, but the programme is still running: later experiments here have already overturned earlier readings more than once, and several pages record exactly that. Expect this section to change as the work moves forward.
Programme two: how learning happens
The Archive as an Instrument
A pooled correlation of -0.679 that dissolved on stratification, and a scaling law that survived. Also the discovery that our archive varied only one of the four settings we thought it had.
Read the recordConsolidation Is a Measurement-Frame Effect
The same 41 runs show a number falling and rising over the same interval, depending only on whether the measuring frame is allowed to move. Nothing contracts.
Read the recordA Task Battery With Designed Headroom
Two of four candidate tasks accepted. The rejections are the informative part: one is all-or-nothing per seed, the other is trivial for a recurrent model.
Read the recordDoes the Effect Survive the Optimizer?
The effect is within 8% across three optimizers, so it is not an artifact of adaptive scaling. Its coincidence with the moment of learning is an artifact of one of them.
Read the recordLatent Knowledge Before Behaviour
The programme's first genuinely early signal, plus the instrument failure that nearly buried it: an under-powered probe produced a confident wrong negative.
Read the recordDoes the Transition Create Features, or Select Them?
A fixed random memory never reaches full accuracy at any width tested, and at a compute-matched budget improves 18.7 times more gradually. The sharp part of learning needs a memory that can change.
Read the recordIs Transition Timing Controllable?
A step, not a dial: any interference costs about 21 steps and far more interference costs barely more. Nothing makes learning arrive sooner.
Read the recordThe Phenomenon Map: Which Architectures Show It
Nine of nine gated recurrent cells show the effect; zero of three others do. The decisive cell is a transformer that learns the task, takes the longest of anything in the grid, and still shows nothing.
Read the recordDoes the Early Signal Generalise?
The lead is 23.4 steps in one design, 14.4 in another, and unmeasurable in a third. The calibration check passes in all three, so the null is real.
Read the recordWhich Weights Carry the Transition?
Motion localises to one gate, necessity to another. No single parameter group is required: the model routes around every freeze, and the finding lives entirely in what each one costs.
Read the recordOne Jump, or Several?
What looks like one sudden jump is four, about sixteen steps apart, ordered by which element of the pattern is being predicted. The order is reproducible at +0.983 across seeds.
Read the recordProgramme one: representation geometry
Wondering if your data is AI-ready?
Start with a free readiness scorecard you run yourself. We never see your data.
Get the free scorecard