Research question

What actually happens at the moment a model learns?

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.

The jump from guessing to competent takes a few dozen steps. What is going on inside the model during those steps?

Where this stands

It builds machinery rather than selecting it, working through the task in a reproducible order and trying a simpler wrong rule on the way. The visible training curve cannot tell you which is happening.

This is where the commercially relevant finding lives: a model handed finished machinery shows the same training curve as one building it from nothing, so the curve alone cannot tell you what you paid for.

Still open · 9 published results bear on this question.

What this does not settle yet

Stated plainly, because the gaps are as much a part of the record as the answers:

  • Why is the first-acquired part the most expensive to remove? The mechanism is unexplained.
  • Everything here is one architecture family on one task family.

How these results fit together

Models on these tasks do not improve gradually. They sit near chance for a long time and then climb steeply, and these nine records are about what is inside that climb. The most useful finding is that it is not one event: what looks like a single jump is several acquisitions in a reproducible order, roughly sixteen steps apart. That ordering is stable enough across seeds to be a real property of the task rather than noise. Two records complicate it in ways worth knowing. First, some of that apparent separation is an artefact of scoring by accuracy, which is a winner-takes-all measurement -- scored in bits instead, the acquisitions are about half as separated, though the *order* is untouched. Second, a model handed finished internal machinery largely skips the reorganisation, while one handed a perfect external shortcut does not, which points at the climb being the construction of machinery rather than the discovery of an answer. The most striking negative here is that two models can have nearly identical improvement curves while one is building something and the other arrived with it. The curve alone cannot tell you which you are looking at.

The results

Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.

One Jump, or Several?

What looks like one sudden jump is four, about sixteen steps apart, ordered by which element of the pattern is being predicted. The order is reproducible at +0.983 across seeds.

Read the record

The Training Curve Cannot Tell You

Give a model finished internal machinery and the reorganisation stops, though its accuracy still jumps. The visible curve looks the same either way.

Read the record

It Tries The Wrong Rule First

The model's answers resemble one simple rule early and a different, correct one later. Its leftover mistakes are systematic rather than random.

Read the record

We Gave It The Answer Anyway

Handing a model a correct shortcut changes nothing about how much it rearranges inside. Handing it finished internal machinery does. The difference says what the rearranging is.

Read the record

The Scoring Made The Steps

Scored a second way with no right-or-wrong line in it, the stages of learning are about half as separated. The order they arrive in is unchanged.

Read the record

Does the Transition Create Features, or Select Them?

A fixed random memory never reaches full accuracy at any width tested, and at a compute-matched budget improves 18.7 times more gradually. The sharp part of learning needs a memory that can change.

Read the record

Learning It Again Is Not Like Learning It

The same amount of learning, six times less internal upheaval. Which means the signature marks learning from scratch, not learning in general.

Read the record

The Habit It Did Not Need

Removing the habit costs almost nothing, so it is not a stage the model needs. What each removal costs turns out to follow the structure of the task itself.

Read the record

A Dial, Not a Switch

Scramble a quarter of the labels, then half, then all of them. The effect shrinks smoothly with each step, and two of our measurement attempts failed first.

Read the record

Back to all research questions

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.