No Other Claim Ran Hot
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.
In plain English
What we asked. Our warm-up results fell apart because they were measured against training that took steps too large for the task. We checked whether any other result in the project had the same problem.
What we found. None had. Every other standing speed-up result was trained at or below its task's best learning rate. Three older results were on tasks where nobody had yet found that best rate, and we have queued a check on one of them.
Why it matters. When a mistake turns up in one place, look for it everywhere. Here the search came back clean, which is worth knowing too.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. An archive audit: no training. The audit script and its kill test were committed before it ran.
Program v2 Bucket A, item A37. Decisive computation: . Output: analysis/hot_rate_audit.py.analysis/hot_rate_audit.json
The question
The data-warm-up thread (A29-A35) found that training at a constant learning rate above a task's tuned rate, with no rate warm-up, damages the first steps -- and anything that softens those steps looks like a speed-up. R18 audited the archive's speed-up claims for the opposite exposure, a rate slower than tuned. Did any other surviving claim run hot?
The audit: every learning-rate literal of 0.008 or more in analysis/*.py and configs/**/*.yaml, with the modules that inherit it by import; R18's standing speed-up claims against their substrate's tuned constant rate (0.006 on J8's delayed copy, O16); and the thread's own two tasks (copy lag 8 tuned at 0.003, A31; modular sum best at 0.002 or below, A32). Kill test, fixed before execution: no surviving claim outside the data-warm-up thread is HOT.
Results
| Claim | Rate | Substrate's tuned rate | Verdict |
|---|---|---|---|
| R18's nine claims on J8's delayed copy (P1, P6, P7, P8, O10, K2, J7, J8, O14) | 0.002-0.005 | 0.006 | not hot |
Q4 (dispatch copy, lag 8) | 0.005 | 0.010 (O19) | not hot |
| R18's three claims on other substrates (P11, R2, R12) | 0.005 | none committed | no tuned rate |
| A18 (modular sum) | 0.005 | 0.002 or below | hot -- under test (A36) |
| A15 and the copy-task thread | 0.01 | 0.003 | hot -- closed (A31, A33, A35) |
The kill test fires: nothing outside the data-warm-up thread ran hot. Every rate literal of 0.008 or more traces to the thread (lag_staircase.py and sum_mixture.py, 20 modules with those that import them), to program v1's low-rank configs (a tiny transformer at 0.01, closed in 2026-08 as "geometry is not load-bearing" and never an efficiency claim), to the smoke-test config, or to an SGD rate (0.2) in I1's probe study, where SGD's scale differs from AdamW's and no speed-up is claimed.
The remaining exposure is the three claims with no tuned rate on their substrate (R18 marked four UNTESTED; Q4's substrate has since been swept by O19). They ran at 0.005, which is below J8's tuned rate but above the copy lag-8 task's (0.003); without a sweep on their own tasks, whether 0.005 is hot there is unknown.
What stands
- A37: kill test fires. No surviving speed-up claim outside the data-warm-up thread ran above its substrate's tuned rate without a warm-up.
- Three claims (P11, R2, R12) have no tuned rate on their own task. A rate sweep there would close both R18's and A37's exposure at once.
Limits
- A correction before publication: the audit's first run classed Q4 as having no tuned rate, missing O19's committed sweep of dispatch copy. The script was corrected in a separate commit and rerun; the kill test's outcome did not change.
- A literal search finds rates set as constants or config values; a rate computed at run time would be missed. The claims list is R18's, which covers speed-up claims to 2026-09-26.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- accuracy
- The fraction of answers a model gets right on questions it was not trained on.
- AdamW
- A widely used training algorithm that adapts how big a step it takes for each individual parameter. The default in most modern AI training.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- learning rate
- How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
- probe
- A small separate model trained to read information out of a bigger model's internals, used as a measuring instrument rather than as a product.
- rank
- How many independent directions a set of numbers really uses. A low-rank structure is one that looks high-dimensional but is actually simple underneath.
- SGD
- Stochastic Gradient Descent. The simplest training algorithm: take a step in the direction the gradient points, every time, with no adaptation.
- transformer
- The architecture behind most modern large language models. It uses attention to look at every part of the input at once, rather than reading in order.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.