Research question

What actually makes training cheaper?

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.

Training a model costs money and time. This is where we test the ideas that are supposed to reduce that cost, and report what they actually save once the cost of running them is counted.

Where this stands

One thing has worked: stopping part of the training early saved about 7% with no loss of quality. Everything else tested has been matched by a simpler or cheaper method -- and in two cases the clever method was only winning because it was quietly being given more.

This is the question the whole project exists to answer, and the one where it is easiest to fool yourself. A saving is only real once you have subtracted what the method cost you, in the same units, and compared it against the simplest thing that could have worked instead.

Still open · 4 published results bear on this question.

What this does not settle yet

Stated plainly, because the gaps are as much a part of the record as the answers:

  • The one real saving came from a fixed schedule, not from measuring anything. We do not yet know of a saving that requires knowing what the model is doing.
  • Nothing here has been tested at a scale where an instrument's cost would amortise, or on anything other than a small model on a synthetic task.
  • Bucket O is mostly aimed at this question: progressive width, function-space accounting, and whether any of it survives a compute-matched control.

How these results fit together

This is the question the project exists to answer, and it is the one where it is easiest to fool yourself, so the standard here is deliberately harsh: a saving counts only once you have subtracted what the method cost you, in the same units, and compared it against the simplest thing that could have worked instead. By that standard one thing has worked. Freezing the expensive recurrent component late in training saved about 7% of the run with no loss of quality. The instructive part is that timing it from the model's own behaviour gave no advantage over picking a fixed step in advance -- the saving is real and the measurement that produced it was not needed. That has now happened three times in different guises, and it has become a standing rule here: find the best offline schedule before inventing anything adaptive. Two attempts failed in a way worth reporting. Choosing training data cleverly barely beat taking whatever came first, even with an unrealistically expensive selector. And a method that looked fastest turned out to be winning because it was quietly being given more than its competitors -- a reminder that most efficiency claims die on their controls, not their ideas.

The results

Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.

We Found a Saving, and It Needed No Instrument

Switching off the most expensive part of the model saves 7% of training for free, if you wait long enough. A calendar times it as well as any detector.

Read the record

Predicting It Without Watching

The free formula only wins when model size and exercise difficulty are kept as two separate numbers. Combined into the one number that correctly describes how abrupt learning is, it misprices when learning happens by more than three times.

Read the record

Being Choosy About Data Costs More Than It Saves

Every selection rule we tried, including one with perfect information at impossible cost, was indistinguishable from taking whatever batch arrived first. The economics never arise because there is nothing being saved.

Read the record

The Method That Looked Fastest Was Cheating

It scored perfectly on the easy half and worse than the control on the hard half, which it chose 9% of the time against the control's 51%. Our speed measurement did not notice.

Read the record

Back to all research questions

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.