Can you tell in advance that a model is about to improve?
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
These models sit at chance for a long time and then get good over a few dozen steps. Can you see that coming before it happens?
Where this stands
Yes, repeatedly and by several independent routes -- a probe on the hidden state, gradient statistics, held-out loss, and which earlier word the prediction leans on. The warnings are real; how far ahead they fire varies a great deal.
An early warning is the most obviously useful thing this project could produce, and it is the claim most often made loosely elsewhere. Every warning here is reported with its cost and its failure rate.
Still open · 15 published results bear on this question.
What this does not settle yet
Stated plainly, because the gaps are as much a part of the record as the answers:
- Does any warning survive on a task a lookup cannot solve? (N18)
- What predicts the LENGTH of the lead? The spread is what makes it unusable.
- Every warning here is coincident-or-leading on one event type; none has been tested on a model that fails to learn at all.
How these results fit together
Fifteen records, and they divide into three groups that are easy to confuse. The first group found warnings that are real: a small probe reading the model's internal state can decode the answer before the model will say it, held-out loss moves before held-out accuracy, and gradient statistics shift around the same time. The second group asked how far ahead those warnings actually fire, and this is where the picture tightened -- the best instrument warns reliably about 45 steps ahead, only about half the time at 25 steps, and almost never at 10. The third group is the one that matters most, and it is mostly negative: several apparent warnings turned out to be artefacts of how we were measuring. An alarm at a fixed step looks like a detector until you move the event and watch it fail to follow. A spectral peak we thought was a signal was not. Combining two warnings cannot help, because a rule that waits for both fires at the later of them. The useful summary is that the warnings are genuine, they are not precise, and most of the ways of making them look better are ways of fooling yourself.
The results
Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.
Latent Knowledge Before Behaviour
The programme's first genuinely early signal, plus the instrument failure that nearly buried it: an under-powered probe produced a confident wrong negative.
Read the recordDoes the Early Signal Generalise?
The lead is 23.4 steps in one design, 14.4 in another, and unmeasurable in a third. The calibration check passes in all three, so the null is real.
Read the recordPredicting When a Model Will Learn
Gradient variance leads the moment of learning in 49 of 49 runs and moves when the moment moves. Two of four statistics are inconclusive for instrument reasons, so the theory behind it is only half confirmed.
Read the recordAn Average Is Not a Guarantee
Two of 24 candidates survive and replicate. The best tracks at slope 1.01, which looks like a fixed warning window, and still fires late in 5 of 31 runs. A controlled comparison shows our earlier 49-of-49 figure was the run population, not the detector.
Read the recordThe Spectral Peak That Was Not a Signal
Eleven of twelve measurements led by up to 74.5 steps. Zero of eleven followed the transition when we moved it 145 steps. The lead was the gap to a fixed early moment.
Read the recordTwo Warnings That Know the Same Thing
Two moderately correlated signals, and combining them never helps. The more useful discovery was how unreliable each one is once you ask it to be precise.
Read the recordThe Warning Was Not a Quirk of One Method
It survives, and under the simplest training method it arrives even earlier. Our control reproduced the published figure to the decimal.
Read the recordWe Let It Learn How to Combine Them
It looked like a win until the control. Almost all of it came from a better cut-off, not from the second signal, and four earlier versions were quietly broken.
Read the recordThe Warning Is a Share of the Run
We asked what makes the warning longer in some runs than others. The answer was that we had been measuring it in the wrong unit all along.
Read the recordDoes the Warning Cry Wolf?
On models with nothing to learn it never fired once. On models that were merely slow it fired more than a third of the time.
Read the recordThe Warning Runs Out of Room
Both our warning signals fail in bigger models, for opposite reasons. One fires too early to be useful; this one cannot fire early enough.
Read the recordOur Early Warning Gets Worse as Models Grow
It does not get noisier at scale. It fires earlier and earlier until it is telling you almost nothing, and it corrects an earlier claim of ours.
Read the recordThe Early Warning We Already Had
First we checked that the sudden jump these models make is real and not an artifact of how we score them. It is. Then we noticed the free early warning we had walked past fifty-four times.
Read the recordLooking In The Right Place First
The longest early warning we have found, and it comes from a ranking rather than a threshold. A solved model also depends heavily on words that cannot help it.
Read the recordThe Check That Could Have Sunk It
The measurement might have been reading the exercise rather than the model. A harder version, plus a deliberately misaligned control, shows it is reading the model.
Read the record