The Number Was Right. The Advice Was Wrong.
We run a research programme on models small enough that every experiment can be rerun on a laptop in minutes. It exists so that what we tell clients rests on a process we can show you rather than on results we assert.
Last week we counted something uncomfortable about our own record: of forty-three findings that were still standing, four had ever been checked on more than one kind of task. Nineteen rested entirely on a single one. Nothing was withdrawn by that count. It simply moved nineteen results from "confirmed" to "unchecked", which is a different thing and had never been written down.
This is what happened when we went back and checked the first of them.
Which one to check first
Not the cheapest. The one that the most other work leans on.
The finding in question is the only one we have with a break-even attached. It answers a practical question about reusing models: if you train a model once and reuse part of it to give later models a head start, how many later models does it have to help before the first one has paid for itself? The answer we published was about five.
Every downstream piece of that work quotes that number. If it is only true of the one task we measured it on, a lot of other writing needs a caveat it does not currently carry.
The number survived
We rebuilt the whole experiment on a different task. Different structure, different thing to learn, about twice as long for a model to figure out at all. Then we scaled the schedule so the comparison was about economics rather than about step counts, and re-ran it.
The best answer on the new task is 5.3 models. On the original it was 4.8. Within ten percent, on exactly the arm the argument depends on. We also required, in code, that the run reproduce an unrelated published measurement before we were allowed to interpret anything. It reproduced exactly.
That is the first of the nineteen re-tested, and it held. Good.
The advice did not
Underneath that number we had written a rule of thumb, in the plain way anybody writes one: the cheapest model that helps at all is the best one to use. On the original task that was simply true. Train the donor model a little longer and it saves the next model slightly more, but it costs a lot more to make, so the economics get steadily worse. Cheapest useful wins.
On the new task the curve has a dip in the middle.
The cheapest model that helps at all is a bad choice there. Following our own rule picks it, and it lands at 11.3 models to break even, which is over the line we had set in advance for the idea being worth doing at all. The actual best choice is a longer-trained model at 5.3. The rule costs about double the answer.
The reason is worth more than the result. Our original experiment started its range at what turned out to be the optimum, and went up from there. We only ever saw one side of a hill and reported it as a slope. The rule was a faithful description of everything we had looked at.
Why this matters outside our lab
Three things travel from this, and none of them is about small models on synthetic tasks.
A number generalizing is not the same as the advice drawn from it generalizing. These get reported together and they are not the same claim. The number is a measurement. The advice is a claim about the shape of a curve away from where you measured, and shape is exactly what a narrow range cannot tell you. When a vendor gives you a benchmark figure and a recommendation in the same breath, those two things need checking separately.
Check whether your range covers the answer. If your best result is at the edge of what you tried, you have not found an optimum. You have found the edge. That is a one-line check nobody runs often enough, ourselves included, and it is the entire reason the original rule was wrong.
Re-test in order of what depends on it. Not in order of what is easy. The nineteen unchecked findings are not equally load-bearing, and the one we started with is quoted by everything else in its thread. Doing the cheap ones first would have felt productive and told us nothing about our exposure.
The thing we did not expect
There was a third result, and it exists only because the new range went below the optimum instead of starting at it.
A model reused by somebody else has to be more fully trained than one reused by the team that made it, before it is worth the same. Early on the gap is large. It closes as the donor model converges, until at the long end the two are indistinguishable.
Our earlier work had found those two cases equivalent and treated it as settled, because it only ever looked at the long end. This does not overturn that result. It bounds it: the equivalence holds for a converged model and not before, which matters quite a lot if you are planning to hand a model around a company.
It also rests on a single donor, which is thin, so it is now a queued experiment rather than a claim. That is on the record too.
What we did not do
We did not correct the original write-up. Its number held, and its advice was stated about the task it measured, which is what an honest record looks like. What changed is the scope, and scope belongs in the new record rather than as a retroactive edit to the old one.
Two of the three predictions we wrote down before running this were right. The third was wrong, and it is the only part anybody should remember.
Read the full technical record
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.