Research practice

The Day Our Results Turned Out to Be Our Own Mistake

Over the past weeks our open research programme accumulated six results with the same shape. Each one found something interesting happening at a particular moment during a model's training, tested whether it still happened in bigger models, and found that it did not. Six for six.

That is the kind of consistency that stops feeling like a series of results and starts feeling like a fact about the world. We wrote it down as a rule: before building anything on a finding of that shape, test it against model size, because it will probably not survive. The rule felt like hard-won discipline.

It was wrong, and the way it was wrong is worth more than the six results were.

What all six had in common

In every one of those tests we made the model bigger and left the task exactly as it was. That sounds like good practice, because you are changing one thing at a time. But a bigger model on the same task is not simply a bigger model. It is a model with less to do.

Our largest models were finishing the task in a third of the time the small ones took. So when we looked for something happening partway through training, there was barely any training left for it to happen in. We were not measuring how these effects behave at larger scale. We were measuring how they behave when a model has far more capacity than the job requires, which is a different question and not the one we were answering.

The check that settles it

The fix is to make the task harder as the model gets bigger, so the model stays under roughly constant strain. We tuned each model size to a difficulty where it took about the same number of steps to learn, then re-ran one of the six.

The effect did not fade at all. Where the original test showed it shrinking to almost nothing at the largest size, the corrected version shows it flat: the same size at the largest model as at the smallest. At the biggest model the two versions differ by a factor of fourteen, and the only difference between them is how hard the task was.

We then re-ran a second of the six, chosen because it was the most consequential: a short window during training where interrupting part of a model does far more damage than interrupting it earlier. We had reported that window vanishing in our largest model, which scoped down a lot of our own work, including the only genuine efficiency result we have. Given a harder task, the window is back and comfortably above our own threshold for mattering.

Two out of two. Four still to check, and we are not assuming they will go the same way.

Why it was so convincing

This is the part we think generalises past our small models.

The six results were persuasive precisely because they agreed with each other. One result fading might be noise. Six fading looks like a law. But they agreed because every one of them was generated by the same procedure, and the procedure had the flaw. A systematic mistake does not announce itself as noise. It announces itself as consistency, which is the thing we are all trained to find reassuring.

There is a general lesson in that for reading anyone's results, including vendors'. A pile of experiments that all point the same way is only as independent as the method behind them. If they share a pipeline, they share its blind spots, and the agreement between them tells you much less than it appears to.

What actually caught it

Not an experiment. We had two papers sitting on a list to read, flagged because they touched results we had already published. One of them contained a sentence about task difficulty relative to model capacity that did not fit our picture, and following that thread led straight to the flaw.

Reading cost no computing time at all and was more valuable than anything we ran that day. We keep a list of cheap items that look low-value next to a real experiment, and this is the second or third time one of them has turned out to be the most important thing on it.

Worth adding: we got the paper partly wrong on first reading too. We wrote our initial summary from abstracts, said so at the top of the page, and then read further and had to correct three claims about what that paper actually says. The caveat did its job. The underlying problem with our own work survived the correction and came out sharper.

What we changed

  • Six pages now carry correction notices. We append them rather than editing the original text, so the record shows what we thought and when we stopped thinking it. Every number in those records still reproduces exactly. What failed was interpretation, not arithmetic.
  • The rule got rewritten, not patched. It was pointing at a real worry with an instrument that manufactured the very effect it reported.
  • Two of our conclusions came back. Results we had scoped down as small-model artefacts are, on this evidence, not limited that way.
  • Four results are still unchecked and are listed as uncertain rather than quietly assumed to be fine.

None of this makes the underlying models less small or the task less synthetic. Those limits are on every page and they have not changed. What changed is that a conclusion we were confident about turned out to be a property of how we ran the test, and we would rather publish that than the tidier version.

{_repo( "experiments/studies/matched-difficulty-v1.md", "The full record, with both ladders and every number", )}

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.