Research practice

The Experiment That Cancelled Our Next Six Months

The most valuable experiment we ran this month took an afternoon and produced nothing usable. It cancelled roughly six months of planned work, which is exactly what it was designed to do.

The plan it cancelled

Models in our research repository learn suddenly. They sit at chance for a while, then acquire the task in a burst of a few dozen steps. We had spent months characterising that moment, and had a natural next question: could you steer it? Make it happen sooner, more reliably, on demand?

An entire branch of the plan depended on that being possible. It needed two things:

  1. Advance warning. Some signal that the moment is coming before it arrives.
  2. A response you can dial. A small intervention producing a small change, a larger one producing a larger change.

We had just found the first. A linear probe gives about 23 steps of warning. That was the encouraging result, and it made the second question worth answering properly.

The test

Interrupt training partway through. Nudge the model by a controlled amount. Let it continue, and see how much later it learns.

To make the comparison mean anything, the replay has to be exactly reproducible: same data in the same order, same internal state, everything. So the control case is a nudge of size zero, which must land in precisely the same place as the original run. Ours does, in every trial.

Bar chart showing near-identical delays for nudges spanning a factor of ten million
How much later the model learned, against the size of the nudge. The leftmost bar is the control: change nothing, and training lands in exactly the same place.

There is no dial

The smallest nudge we tried alters the model by one part in a hundred million. It delays learning by about 21 steps.

The largest nudge is ten million times bigger. It delays learning by about 24 steps.

Ten million times the cause, fourteen percent more effect. And every single measurement was a delay. Nothing we did made the model learn sooner.

In plain English

Imagine a light switch you were hoping was a dimmer. You want to set the brightness precisely, so you try nudging it very gently, and the light goes off. You shove it hard, and the light goes off by the same amount. There is no in between, and no direction that makes it brighter.

You can interfere with the timing of learning. You cannot tune it. Those are different things, and only one of them supports building a controller.

So we cancelled it

Having advance warning without a usable response is like a weather forecast for a planet you cannot travel to. Genuinely interesting, not actionable in the way the plan required.

The planned work is now closed in the repository, marked with the result that closed it. That is the entire purpose of designing an experiment that can kill an idea, and running it before the expensive work rather than after.

The near miss

One detail, because it nearly went the other way.

Our first version of this test measured progress every five steps, which is what the rest of the project uses. It reported the same delay at every nudge size, and we could have stopped there.

But that reading is ambiguous. A response that genuinely ignores the size of the nudge looks exactly like a real response measured on a grid too coarse to see it. A stopwatch that only ticks in five-second intervals will report the same time for every runner in a close race.

Those two readings have opposite consequences, and this one closed a large chunk of our plan, so we re-ran the whole thing measuring every single step. The answer held. Had it not, we would have cancelled the work for the wrong reason.

That is the third time this year a result in our repository turned on the resolution or the strength of the measuring tool rather than on the system being measured. It is becoming the most common way we are wrong.

Why we are writing about a dead end

Because the alternative is what usually happens. Nobody publishes the plan that got quietly dropped, so the same idea gets re-attempted, and the cost of finding out is paid again.

Two things transfer from this to anyone evaluating AI work, ours included:

  • Ask what would have falsified it. A programme of work with no experiment capable of ending it is not a research plan, it is a budget. Ours had one, it cost an afternoon, and it fired.
  • Ask whether the measurement could see the answer. A clean result from an instrument that was never checked against a known case is a statement about the instrument. We have now been caught by that twice.

The warning signal survives, and it is still useful: knowing that a model is about to learn something is worth having even if you cannot make it happen sooner. It just supports a different, smaller set of decisions than the ones we had planned.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.