How do we know our own results are real?
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
Most of what goes wrong in this kind of work is not the experiment; it is the comparison. These are the times our own checks overturned our own headline.
Where this stands
Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
If you are evaluating whether someone would catch a problem in your models, this is the group to read. It is the most transferable output the project has.
Still open · 10 published results bear on this question.
What this does not settle yet
Stated plainly, because the gaps are as much a part of the record as the answers:
- The programme still has no unlearnable-data control anywhere except one record (I6).
How these results fit together
Ten records that are all about this project checking its own work, and they are the most transferable thing here. Several overturned our own headlines. The largest is that a pattern we had described as the model expanding and then consolidating is substantially a *measurement* effect: we were re-fitting the coordinate frame at every step, and the apparent contraction is the frame rotating rather than the model unwinding. Any quantity compared across training steps now has to say which frame it was measured in. A companion record does the opposite service -- certain quantities depend only on eigenvalues, which are unchanged by rotation, so that critique cannot reach them, and one of those still peaks on the event. We also audited how many of our own results are close calls, found that eight runs is not enough for several of them, and published that. The habit that produced most of these: re-read the whole record after every new result, not just the item in front of you. The number of times the mistake was in the *analysis* rather than the experiment is the single most useful statistic on this site.
The results
Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.
The Check We Should Have Run First
We gave a model the same inputs with random answers, so there was nothing to learn. The effect we have been reporting all along vanished, which is what we hoped.
Read the recordEight Runs Is Not Enough
The usual statistical test said we had found something. The range of values the data was consistent with said we had not. Repeating it on three times as many runs settled it, and the answer was yes after all.
Read the recordHow Many of Our Results Are Close Calls
Sixteen of them are within one error bar of their own verdict. We are publishing the list rather than quietly rerunning the convenient half of it.
Read the recordWe Read Our Own Warning List
Only two were things a conclusion depends on, and both had already been rechecked. The tool ranks careful studies as the shakiest, which is backwards.
Read the recordConsolidation Is a Measurement-Frame Effect
The same 41 runs show a number falling and rising over the same interval, depending only on whether the measuring frame is allowed to move. Nothing contracts.
Read the recordThe One That Was Not A Ruler
A measure that cannot depend on which directions you call important rises to a peak exactly when the model learns, then falls back. So the change is real.
Read the recordThe Fix That Fixed Nothing
The one dial we had made the task fainter rather than deeper, and faint is hard for everyone. We learned what the next attempt actually has to do.
Read the recordWe Tested Someone Else's Prediction
Cutting the route made no difference at all. That rules out one of the three explanations the paper offers, and leaves two standing.
Read the recordThe Experiment We Could Not Run
The control worked. The comparison arm never produced anything to compare, and reporting its result would have looked like a finding when it was evidence the task was wrong.
Read the recordThe Archive as an Instrument
A pooled correlation of -0.679 that dissolved on stratification, and a scaling law that survived. Also the discovery that our archive varied only one of the four settings we thought it had.
Read the record