The Same Confound In Six Results
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
Part of a bigger question: How do we know our own results are real? – Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
In plain English
What we asked. This one is not an experiment. We had two research papers sitting on our list to read, flagged because they touched results we had already published. Reading them turned out to matter more than most of the experiments we ran that day, because one of them contains a sentence that undermines six of our own conclusions.
What we found. Six times now we have found something that happens at a particular moment during training, tested it at larger model sizes, and watched it disappear. We had been reading that as a warning about ourselves: these are small-model effects and probably will not matter in practice. The paper points out a different possibility. In every one of those tests we made the model bigger while leaving the task exactly as it was, so the model had more and more capacity going spare. The paper reports that these kinds of effects survive at large scale when the task stays hard relative to the model, and vanish when it does not. That fits all six of our results, and it is a completely different conclusion: not that the effects are too small to matter, but that they show up when a model is working near its limits, which is where real training actually happens.
Why it matters. We do not know yet which reading is right, and the honest position is that six of our published negatives are now uncertain rather than wrong. The experiment that would settle it makes the task harder as the model gets bigger, so the strain stays constant, and it is now the most valuable thing on our list. Two smaller notes. This same reading confirmed one of our positive findings independently at models roughly a thousand times larger than ours, which is the first outside corroboration we have had. And a second paper that looked like it contradicted us turned out not to, once read carefully: its claim that emergence is random means random around a trend that still depends on model size, which is what we found too. Both of those were settled by reading rather than spending a day of compute.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Negative results are kept. Roughly half of what is published here says an idea did not work, including several of our own. Those pages are not failures, they are the output. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. No experiment was run. This is a reading record. No GPU, no CPU, no cost. It is based on abstracts and search summaries, not full papers. Every claim attributed to either paper below should be checked against the paper before anything is built on it. That limitation is the reason the item was filed as "resolve by reading, not running", and it is not fully discharged.
CORRECTED 2026-09-01, hours after publication, by reading the paper further. The record below was written from abstracts and said so; that caveat did its job. Three things are wrong or overstated, and one is sharpened. The paper does not claim the patterns persist at scale. Its own words: "Whether the geometric anatomy persists at frontier scale (>7B) remains open", and its largest model is 85M, which it calls small by current standards. The summary this record was built from put that far more confidently than the paper does. The easy-task mechanism is not capacity slack. It is resolution: "Easy tasks emerge during the universal collapse window, making any initialization-phase geometric event appear simultaneous with emergence at larger scales." The precursor is not destroyed, it becomes impossible to separate from the event. "Scale-invariant" in its title means the opposite of what this record implies. It refers toRankMe ~ 2.0holding across a 210x parameter range with a coefficient of variation of0.08-- a measure that is constant across scale, not one that fades. The confound in our six results survives, and is better supported than before. As our models outgrow the fixed task, the transition arrives sooner and faster: a control run transitions at130.0at width 24 and at58.8at width 96. O7's analysis window is +/-30 steps, which is over half the entire pre-transition period at width 96. So a fixed-width window cannot resolve at width 96 what it resolves at width 24, and "the effect faded" and "the instrument ran out of room" are not separated by any sweep we have run. That is a resolution confound rather than a capacity-slack one, it is sharper, and it makes P11 more necessary rather than less.
Program v2 Bucket P, item P4. Two 2026 papers were logged in Bucket P as bearing on published records here. This is what they say, and what it does to us.
Paper one against D6 and N9: they agree, as N9 predicted
The Geometric Anatomy of Capability Acquisition in Transformers (arXiv 2602.15997) tracks geometric measures across six transformer sizes from 405K to 151M parameters, eight algorithmic tasks, and three Pythia models from 160M to 2.8B.
Its headline geometric measure is RankMe, an entropy of the normalised singular values. That is a function of singular values alone, and therefore invariant under rotation, so D6's frame critique cannot reach it, exactly as N9 argued for the same class of quantity. This is agreement, not conflict, and it is the outcome N9 predicted for a rotation-invariant measure.
It also reports that "hidden states already contain task-relevant information before the model can act on it", which is F2's finding, independently, at 160M–2.8B. That is the first external corroboration of one of this programme's positives at a scale three orders of magnitude above ours.
Paper two against N6 and E4: also agreement, and the item was right that reading would settle it
Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns (arXiv 2606.25010) reports that capabilities "arise stochastically throughout training, with larger models acquiring them earlier on average."
So "random" means stochastic around a scale-dependent trend, not instead of one. That is compatible with N6's +12.11 sharpness per doubling of width and E4's width dependence beating seed noise 7.4x: a systematic mean with a noisy residual. There is no disagreement to resolve, and it took reading rather than running to establish that, which is what the item said it would.
The part that matters: a confound in six of our own results
Paper one's stated condition is the important sentence:
geometric patterns in small models can persist at larger scale when the task remains difficult relative to model capacity, and easy tasks show no detectable precursor, because learning happens too rapidly.
Every width sweep in this programme holds the task fixed while width grows. So difficulty relative to capacity falls across the sweep, and the largest width is always the point of greatest capacity slack. That is precisely the regime this paper reports the precursor disappearing in.
Six results here fade with width, and this session added a standing rule off the back of them:
- the critical-period window, which had vanished at the largest size tested
- F2's decode-probe lead, which shrinks as models grow
- two K1 gradient alarms
- P9, which closed P8's donor head start
- O7, which closed the weight-spectrum signal, hours ago
All six share the confound. None of them separates "the phenomenon fades with scale" from "the phenomenon fades as the model outgrows the task". Those are very different claims: the first says this programme studies a small-model artefact, the second says it studies what happens when capacity is tight, which is the interesting regime and the one real training operates in.
Bucket M already said this in other words, these tasks are easy and these models are oversized for them, and the connection was not made until an outside paper stated the condition explicitly.
What this does to the standing rule written today
The rule says: width-sweep a timing-shaped result before building on it. It stands, because a sweep that closes is still a warning. But its interpretation was wrong in this record's view, and both the rule and the six records now need the qualification: a width sweep at fixed task difficulty confounds scale with capacity slack, and the sweep that would separate them holds difficulty relative to capacity constant instead.
That is filed as P11, and it is now the most valuable open item here: it does not add a phenomenon, it decides whether six published negatives mean what they say.
Limits
- Abstracts, not papers. The strongest claim above, the difficulty-relative-to-capacity condition, rests on one sentence of summary. It should be read in full before P11 is designed, and if it turns out to be mis-summarised, this record is wrong and P11's premise goes with it.
- Whether RankMe is computed against a re-fitted basis anywhere in that paper is not established here; the argument above is about the quantity's definition, not about their implementation.
- No experiment was run. Nothing here is evidence about our models; it is a re-reading of our own results in the light of somebody else's stated condition.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- confound
- A second explanation you did not control for. If bigger models both learn faster and score higher, then 'fast learners score higher' may be entirely about size and not about speed.
- critical period
- A stretch of time during which something has to happen for development to proceed normally. Borrowed from biology, where it describes windows in which a young brain must receive certain input.
- gradient
- The direction and amount by which each of a model's internal numbers should change to do slightly better. Training is repeatedly following it.
- hidden state
- The model's working memory: the internal numbers it carries from one step of a sequence to the next.
- parameters
- The adjustable numbers inside a model. Training is the process of setting them. Model size is usually quoted as a count of these.
- probe
- A small separate model trained to read information out of a bigger model's internals, used as a measuring instrument rather than as a product.
- residual
- How far the data sits from a proposed description of it, after the description has been fitted as well as it can be. A small residual means the description accounts for what was measured; a large one means something is missing.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
- sharpness
- How abrupt a model's improvement is. We measure the biggest change over any short stretch of training as a share of the run's total change, scaled so that a model improving at a perfectly steady rate would score 1. A score of 20 means the fastest stretch was twenty times steeper than a straight line.
- singular value
- A measure of how much a mathematical operation stretches things along one particular direction. The largest one describes the direction of strongest effect.
- transformer
- The architecture behind most modern large language models. It uses attention to look at every part of the input at once, rather than reading in order.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the experiments that went against us, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.