Which parts of a model actually matter?
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.
When a model learns, some parts change a lot. Are those the parts that matter? And can you find the parts you cannot afford to disturb?
Where this stands
The parts that move most are not the parts that matter, and no single component is required -- the model routes around every freeze. What a part is worth shows up only when you remove it.
Almost every interpretability claim is about where something is. These records are mostly about why "where it moves" is the wrong question.
Still open · 15 published results bear on this question.
What this does not settle yet
Stated plainly, because the gaps are as much a part of the record as the answers:
- Motion, predictive power and necessity have each been measured; no record has all three on the same quantity at once.
How these results fit together
Fourteen records, and one lesson runs through nearly all of them: **the part that changes most is repeatedly not the part that matters most.** We measured which components move during the jump, then removed each one to see what actually broke, and the two rankings kept disagreeing. A compartment can look loud because of how the architecture is shaped rather than because it is doing the work. Models also route around the freezing of any individual parameter group, which makes "this part is essential" a hard claim to earn. Against that, one result points the other way and is worth the attention: removing the single strongest direction in the update did far more damage than removing many ordinary directions carrying three times as much of the force. So importance is concentrated -- just not in the places the architecture diagram suggests. The sharpest negative here is that this concentrated direction does not travel: transplanted into an independently initialised model it delivers a few per cent of the receiving model's own effect, about what a randomly chosen direction delivers. Three of these records still disagree with each other about self-connections, which is recorded rather than smoothed over.
The results
Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.
Which Weights Carry the Transition?
Motion localises to one gate, necessity to another. No single parameter group is required: the model routes around every freeze, and the finding lives entirely in what each one costs.
Read the recordThe Strongest Part of the Push
Removing 14% of the training force from the strongest direction costs 35 steps. Removing 51% from ordinary directions costs 16. Energy is not the operative variable.
Read the recordNo Part of the Model Is Special
Compartments differ from each other by no more than two repeats of the same compartment differ. Cutting the connections did make the model learn later and score worse, so they were doing real work.
Read the recordDepth Does Not Add Stages
The layers are 10 steps apart when each takes 60 steps to change, so they move together. The reorganisation is not uniform though: it nearly doubles from the first layer to the second.
Read the recordThe Layer That Changes Most Is Not the One That Matters
One layer has a real window; the other costs twice as much to interrupt and does not care when. A separate result found the reorganisation concentrates in the layer without the window.
Read the recordThe Loudest Part Is Not the Useful Part
The part of the signal that matters most when you remove it is not the part that tells you most when you watch it. We had assumed those were the same property.
Read the recordIt Is Not How Much You Remove
One direction, 14% of the signal, 61% of the damage. The same direction removed at a random earlier moment takes away three times as much and costs a third as much.
Read the recordHow You Cut a Model Matters, Not Just How Much
Same number of connections removed, clearly different outcomes. Half the advantage turned out to be one specific thing we nearly failed to test for.
Read the recordThe Lean That Carries Nothing
Where something sits and what it does turn out to be different questions. Removing the part we suspected mattered costs almost nothing.
Read the recordThe Number That Did Not Survive Its Check
A striking, memorable, reproducible number that turned out to describe how these models are built rather than what makes them work.
Read the recordThe Direction That Does Not Travel
Whatever makes this direction matter stays inside the model that grew it. Our first answer was the opposite, and a missing minus sign was why.
Read the recordWe Ran It Six Times, Then Twenty-Four
A pattern we spotted in six runs was not there. An effect we ruled out in six runs was. Both corrections came from the same eighteen extra runs.
Read the recordThe Third Bar Changed Everything
Removing the thing we were testing changed the number. So did simply making the model bigger while leaving that thing in place. The two are indistinguishable.
Read the recordThe Waste That Is Not Where You Look
The three buckets the plan called for showed a clean, strong pattern with non-overlapping error bars. The per-step curve shows the pattern is entirely an artifact of the buckets sitting at different times.
Read the recordWhat a Unit of Training Effort Buys
We imported a published way of measuring how much a training step actually changes what a model does, and checked it against an exact calculation. The unevenness it finds is real and large, and both of our controls say training did not cause it.
Read the record