Research question

What decides when a model learns, and can you change it?

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.

Two models given the same task learn at different moments. What sets the timing, and can you move it deliberately?

Where this stands

Settings dominate, data barely matters, and there is a brief window before the jump in which interrupting the model is unusually costly. Timing can be delayed but not brought forward.

A critical window that closes as models grow is the kind of thing that matters when scheduling real training, and it is measurable here in a way it is not at scale.

Still open · 8 published results bear on this question.

What this does not settle yet

Stated plainly, because the gaps are as much a part of the record as the answers:

  • Nothing found so far makes the transition arrive EARLIER. Every lever delays it.
  • The critical window has been measured on one component; whether other components have their own windows is untested.

How these results fit together

Eight records, and they line up into a clear and slightly deflating conclusion: the timing of the jump is mostly decided by ordinary things, and it is not steerable. Ordinary training settings -- learning rate, batch size -- move the moment by hundreds of steps. Changing the entire training dataset moves it by about seven, and changing the order of the data by about three. So the exotic explanations lose to the boring ones by two orders of magnitude. We then tried to move the moment deliberately, and found the response to a nudge is discontinuous: tiny and enormous perturbations both delay learning by about the same amount, and nothing we tried made it arrive sooner. That closed a whole line of work about steering training, and it is why this project stopped trying to build a controller. Alongside that sits a genuinely causal result -- there is a window of about forty steps just before the jump where interfering with the model's memory delays it much more than interfering earlier or later. The window is real, but it narrows as models get bigger and had vanished at the largest size we tested, which is a reason to be careful about building anything on it.

The results

Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.

The Forty Steps That Matter

Freezing the model's memory just before it learns costs 10.6 steps more than freezing it earlier. Forty steps after, the same freeze costs 0.1 steps. A peak, not a ramp.

Read the record

The Window Closes as Models Grow

The effect shrinks steadily with size and vanishes at the largest we tested. Our first three attempts to measure it were wrong, and one check caught all three.

Read the record

We Had Been Calling It the Right Thing

The effect appears in both gated designs and not in the ungated one, which learns the task faster and better. The name held.

Read the record

Is Transition Timing Controllable?

A step, not a dial: any interference costs about 21 steps and far more interference costs barely more. Nothing makes learning arrive sooner.

Read the record

What Actually Decides When a Model Learns

Data is worth about 2.5% of what the settings are worth. The ratio test we built to decide this landed at 0.49 against a 0.50 threshold, and its confidence interval reached 1.58, so it settled nothing.

Read the record

Is Learning a Lucky Accident?

Noise moves the transition 1.70x later, monotonically, and doubles its width. The batch-size sweep that seemed to show the opposite was measuring update count.

Read the record

Does the Effect Survive the Optimizer?

The effect is within 8% across three optimizers, so it is not an artifact of adaptive scaling. Its coincidence with the moment of learning is an artifact of one of them.

Read the record

The Fix That Sharpened It

Cycling the learning rate makes an interruption at an arbitrary moment much cheaper and one at the critical moment no cheaper, so the gap between them widens.

Read the record

Back to all research questions

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.