Open research: how AI models actually learn

We run a public research programme on how AI models actually learn, using models small enough that anyone can reproduce every number on a laptop in minutes. We publish the results and the data behind them here, including the experiments that went against what we expected.

These pages are the records: the design, the numbers, the controls and the limits. Most of them are about a measurement that misled us before we caught it, which is the same class of failure that makes vendor benchmarks unreliable. If you would rather have the story than the record, the blog carries a plain-language version of most of these, written for a reader who has never seen a training curve.

Ongoing research. These are experimental results from active work, not settled conclusions. The numbers are what we measured and the methods are described so you can judge them, but the programme is still running: later experiments here have already overturned earlier readings more than once, and several pages record exactly that. Expect this section to change as the work moves forward.

Programme two: how learning happens

Research record

The Archive as an Instrument

A pooled correlation of -0.679 that dissolved on stratification, and a scaling law that survived. Also the discovery that our archive varied only one of the four settings we thought it had.

Read the record
Research record

Consolidation Is a Measurement-Frame Effect

The same 41 runs show a number falling and rising over the same interval, depending only on whether the measuring frame is allowed to move. Nothing contracts.

Read the record
Research record

A Task Battery With Designed Headroom

Two of four candidate tasks accepted. The rejections are the informative part: one is all-or-nothing per seed, the other is trivial for a recurrent model.

Read the record
Research record

Does the Effect Survive the Optimizer?

The effect is within 8% across three optimizers, so it is not an artifact of adaptive scaling. Its coincidence with the moment of learning is an artifact of one of them.

Read the record
Research record

Latent Knowledge Before Behaviour

The programme's first genuinely early signal, plus the instrument failure that nearly buried it: an under-powered probe produced a confident wrong negative.

Read the record
Research record

Does the Transition Create Features, or Select Them?

A fixed random memory never reaches full accuracy at any width tested, and at a compute-matched budget improves 18.7 times more gradually. The sharp part of learning needs a memory that can change.

Read the record
Research record

Is Transition Timing Controllable?

A step, not a dial: any interference costs about 21 steps and far more interference costs barely more. Nothing makes learning arrive sooner.

Read the record
Research record

The Phenomenon Map: Which Architectures Show It

Nine of nine gated recurrent cells show the effect; zero of three others do. The decisive cell is a transformer that learns the task, takes the longest of anything in the grid, and still shows nothing.

Read the record
Research record

Does the Early Signal Generalise?

The lead is 23.4 steps in one design, 14.4 in another, and unmeasurable in a third. The calibration check passes in all three, so the null is real.

Read the record
Research record

Which Weights Carry the Transition?

Motion localises to one gate, necessity to another. No single parameter group is required: the model routes around every freeze, and the finding lives entirely in what each one costs.

Read the record
Research record

One Jump, or Several?

What looks like one sudden jump is four, about sixteen steps apart, ordered by which element of the pattern is being predicted. The order is reproducible at +0.983 across seeds.

Read the record

Programme one: representation geometry

Research record

Representation Geometry Is Not Load-Bearing

A whole programme of work, written up. We could remove 97% of a model's internal spread for under a tenth of a percentage point of accuracy. Internal coordinates are not where the difficulty lives.

Read the record

Wondering if your data is AI-ready?

Start with a free readiness scorecard you run yourself. We never see your data.

Get the free scorecard