Research practice

Our Two Findings Turned Out to Be One

Our public research repository had two live findings. This is the story of discovering they were the same one.

The two

One. When these models learn a task, their internal state reorganises in a distinctive way at that moment.

Two. The right answer becomes readable from that internal state about twenty steps before the model can produce it. That is an early warning, and it is the only result we have with an obvious practical use.

These are different things. The first is about how information is arranged. The second is about when it arrives relative to when it gets used. There was no particular reason they had to be connected.

Checking the scope

We had recently found that the first result holds for a family of model designs and not for others, including not for transformers, which is what most modern AI is built from. So the natural question was whether the early warning had the same boundaries or different ones.

Bars showing advance warning by model design, with the third bar inside an unmeasurable band
How many training steps of advance warning each design gives, on a task all three can genuinely learn. The shaded strip at the bottom is the range too small for the measurement to resolve.

Same boundaries. Two designs give real warning, of about 23 and 14 steps. The third gives 1.3 steps, which is inside the range our measurement cannot resolve, in every single run.

So they are one finding, not two, and the useful one is narrower than we had been describing it. Not "models signal before they learn". "Models of a particular design family signal before they learn."

In plain English

Suppose you notice two things about a car: it makes a particular noise before it stalls, and its fuel gauge behaves oddly. You treat them as separate observations. Then you test other cars and find both things appear together in one make and are both absent in another. They were never two observations. They were two symptoms of one design choice.

The check that makes the negative worth trusting

We have been caught before by a negative result that was really a broken measurement. In an earlier version of this same experiment our measuring tool was too weak, and it produced a clean, confident, wrong answer.

So this experiment carries a built-in check. At the end of training the tool should be able to match the model's own performance, because it has strictly more freedom to fit. If it falls short, it is under-powered and any negative it produces is about the tool.

That check passes in all three designs, including the one that shows no warning. The tool works there. It simply finds nothing to report.

Without that check, the honest conclusion would have been "we do not know".

A distinction that reversed the answer

One more detail, because it is the kind of thing that quietly determines results.

In the transformer, our readout is consistently a little better than the model's own output, by about the same margin as in the designs that do show a warning. If we had measured "how big is the gap", we would have reported that all three behave alike.

But the gap is not a head start. The two curves rise together, so the readout is slightly better at every point in time rather than arriving earlier. A constant offset is not an early warning: knowing something a bit more accurately at the same moment tells you nothing about the future.

Measuring when each curve crosses a threshold, rather than how far apart they sit, is what separated those. The two framings give opposite conclusions from identical data.

And one thing we did not go looking for

The design that originally showed the 23-step warning showed it again here, at 23.4 steps, on a different task, with a different vocabulary size, a different learning rate, and different random seeds.

We were not testing that. It is the most valuable thing in the experiment: an independent replication of our main positive result, obtained as a side effect of asking a different question.

The takeaway

Two results that look independent may be one result observed twice, and you find out by varying something neither of them mentions. Here it was the model design; in other work it might be the dataset, the language, or the deployment setting.

The practical version for anyone assessing AI claims: a finding's scope is not what the finding says, it is what was varied while checking it. Ours said something about models. It turned out to be about one family of models, and we only know because we built the others.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.