Research question

How do we know our own results are real?

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.

Most of what goes wrong in this kind of work is not the experiment; it is the comparison. These are the times our own checks overturned our own headline.

Where this stands

Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.

If you are evaluating whether someone would catch a problem in your models, this is the group to read. It is the most transferable output the project has.

Still open · 10 published results bear on this question.

What this does not settle yet

Stated plainly, because the gaps are as much a part of the record as the answers:

  • The programme still has no unlearnable-data control anywhere except one record (I6).

How these results fit together

Ten records that are all about this project checking its own work, and they are the most transferable thing here. Several overturned our own headlines. The largest is that a pattern we had described as the model expanding and then consolidating is substantially a *measurement* effect: we were re-fitting the coordinate frame at every step, and the apparent contraction is the frame rotating rather than the model unwinding. Any quantity compared across training steps now has to say which frame it was measured in. A companion record does the opposite service -- certain quantities depend only on eigenvalues, which are unchanged by rotation, so that critique cannot reach them, and one of those still peaks on the event. We also audited how many of our own results are close calls, found that eight runs is not enough for several of them, and published that. The habit that produced most of these: re-read the whole record after every new result, not just the item in front of you. The number of times the mistake was in the *analysis* rather than the experiment is the single most useful statistic on this site.

The results

Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.

The Check We Should Have Run First

We gave a model the same inputs with random answers, so there was nothing to learn. The effect we have been reporting all along vanished, which is what we hoped.

Read the record

Eight Runs Is Not Enough

The usual statistical test said we had found something. The range of values the data was consistent with said we had not. Repeating it on three times as many runs settled it, and the answer was yes after all.

Read the record

How Many of Our Results Are Close Calls

Sixteen of them are within one error bar of their own verdict. We are publishing the list rather than quietly rerunning the convenient half of it.

Read the record

We Read Our Own Warning List

Only two were things a conclusion depends on, and both had already been rechecked. The tool ranks careful studies as the shakiest, which is backwards.

Read the record

Consolidation Is a Measurement-Frame Effect

The same 41 runs show a number falling and rising over the same interval, depending only on whether the measuring frame is allowed to move. Nothing contracts.

Read the record

The One That Was Not A Ruler

A measure that cannot depend on which directions you call important rises to a peak exactly when the model learns, then falls back. So the change is real.

Read the record

The Fix That Fixed Nothing

The one dial we had made the task fainter rather than deeper, and faint is hard for everyone. We learned what the next attempt actually has to do.

Read the record

We Tested Someone Else's Prediction

Cutting the route made no difference at all. That rules out one of the three explanations the paper offers, and leaves two standing.

Read the record

The Experiment We Could Not Run

The control worked. The comparison arm never produced anything to compare, and reporting its result would have looked like a finding when it was evidence the task was wrong.

Read the record

The Archive as an Instrument

A pooled correlation of -0.679 that dissolved on stratification, and a scaling law that survived. Also the discovery that our archive varied only one of the four settings we thought it had.

Read the record

Back to all research questions

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.