Research record 39 of 40

We Read Our Own Warning List

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.

In plain English

What we asked. We recently had a result overturned simply by running it more times, so we built a tool to find every other result of ours that was decided by a similarly small margin. It flagged sixteen. That sounded alarming, and the tool itself warned that it could not tell the difference between a headline finding and a minor number reported along the way. So we read all sixteen.

What we found. Only two were things a conclusion actually depends on, and we had already rechecked both by running them three times over. Five of them were our own comparison runs: experiments we set up deliberately so that nothing should happen in them. A comparison like that produces a number sitting on zero, and sitting on zero is exactly what our tool calls a close call. So the more carefully a study is controlled, the shakier the tool makes it look. Its ranking of which studies to worry about is close to backwards.

Why it matters. One study was flagged as our second most worrying on the strength of five numbers, none of which its conclusion uses; that conclusion rests on a comparison the tool never examined. And the two real cases have settled into a different kind of statement than we expected. After running them three times over they are still close calls, which no longer means we looked at them too briefly. It means the effects are genuinely small, and measured accurately. Those are different problems needing different responses, and the original tool could not tell them apart. We are keeping the tool, because it did find the one result worth rechecking, but we are recording plainly that we would have believed its ranking if we had not read the list.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Negative results are kept. Roughly half of what is published here says an idea did not work, including several of our own. Those pages are not failures, they are the output. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. 48 training runs (E7 at 24 seeds) plus an archive-only re-read, no GPU, no cost.

Program v2 Bucket J, item J2. Decisive computation: analysis/power_audit.py (the ROLES and RESAMPLED tables) and analysis/relearning_event.py at 24 seeds. Outputs: analysis/power_audit.json, analysis/relearning_event.json.

The question

J1 found 16 published intervals within one half-width of the line that decides them and said, in its own Limits, that it "cannot tell a headline from an intermediate value" and that reading them is the next step. This is that step, plus the re-run of the worst case.

Most of what our audit flagged was our own controls working correctly
Most of what our audit flagged was our own controls working correctly. We built a tool to find results decided by a very small margin, and it flagged sixteen of ours. Then we read all sixteen and sorted them by what they actually are. Only the orange bar is a result any conclusion depends on. The green bar is comparison runs we deliberately set up so that nothing should happen in them. The tool was mostly finding our own controls. A comparison run designed to show nothing produces a number sitting on zero, and sitting on zero is exactly what our tool calls fragile, so the more carefully a study is controlled the shakier it looks. Only two flagged results were things a conclusion rests on, and we had already rechecked both by running them three times over. The tool still earned its keep: it found the one worth rechecking. But its ranking of which studies look shaky is close to backwards, and we would have believed it if we had not read the list.

Kill test: every fragile headline survives at 24 seeds.

Part one: E7, the most fragile of all, survives

J1 ranked E7's frozen-frame paired endpoint first, clearing zero by 0.060 half-widths against I3's 0.042 before that record reversed. Re-run at 24 seeds:

Endpoint8 seeds24 seedsMargin
swap, moving frame+0.0322 [+0.0143, +0.0500]+0.0329 [+0.0255, +0.0403]0.803.45
swap, frozen frame+0.0158 [+0.0009, +0.0307]+0.0145 [+0.0072, +0.0219]0.060.99

24 of 24 swapped runs produce a second learning event; 0 of 24 controls do. Nothing in E7's verdict changes, and its "immaterial" half is now more secure than before.

The contrast with I3 is diagnosable in advance, and is the reusable part:

I3 (reversed)E7 (survived)
point estimate on 3x the seedsmoved a lot (1.66x1.02x)barely moved (+0.0158+0.0145)
what was actually wrongthe discovery sample was biased: a pattern noticed in the data, then tested on that same datanothing; it was merely imprecise

A fragile interval with a stable point estimate is underpowered. One found by noticing a pattern is a different and worse problem. Re-run the ones that were found by looking, first.

Part two: reading the sixteen changes J1's picture

Every flagged interval was read against its own record's stated endpoint, and the classification is recorded in the script as a table rather than in prose, so it can be argued with:

RoleCountWhat it is
headline2the endpoint a verdict actually rests on
control5an arm built to show nothing
component2an unpaired half of a paired endpoint
intermediate6a descriptive value no verdict rests on

Only two of the flagged intervals are endpoints anything rests on, I3's paired difference and E7's frozen-frame endpoint, and both have already been re-run at 24 seeds.

The scan systematically over-flags careful work

This is the finding, and it inverts J1's file-level ranking.

A control is built to be flat. A flat interval sits on zero. So a good control is always flagged as fragile, and a record with many careful controls is ranked as the most precarious in the programme. Five of the sixteen are the standing random-time-matched controls this programme requires in every geometry experiment; one is I4's overfitting diagnostic, where straddling zero is the desired outcome.

E8 is the clearest case. J1 ranked modular_state second with five fragile intervals, all of them per_block means. E8's verdict is a spread ratio, between-block spread over within-block seed spread, 1.02, plus an argmax consistency count. No per-block interval decides anything. Five of J1's sixteen were pointing at a record whose endpoint the scan never looked at.

And what remains is a statement about effect size, not power

Both surviving headlines are still within one half-width at 24 seeds (0.75 and 0.99). After quadrupling a sample, that is no longer a statement about how hard we looked. These are small effects, precisely estimated, and going to 96 seeds would tighten the interval without changing what it says.

That distinction, underpowered against small, is the one J1 could not make and this one can.

Verdict

  • The kill test does not fire in the way J1 implied it would. Of 16 flagged intervals, 2 are headlines, both already re-run, both settled.
  • E7 survives at 24 seeds and its verdict is unchanged.
  • J1's file-level ranking is misleading and should not be used as a list of suspect records. It ranks by proximity to a boundary, which rewards having controls.
  • The audit remains worth having: it caught E7 and forced the reading, but its output is a list of intervals to classify, never a list of records to distrust.
  • The ROLES table is now the durable artifact, not the ranking: it is the record of which quantity each verdict actually rests on, which nothing in this programme had written down.

Limits

  • The role classification is a judgement, not a measurement. It was made by reading each record's stated endpoint, and it is committed as a table so it can be disputed; a different reader could reasonably move a component to a headline.
  • Only E7 was re-run. I3 was already at 24 seeds. Nothing else was, because nothing else needed it once the roles were read, which is itself a claim resting on the classification above.
  • The 1/sqrt(n) argument remains a rule of thumb, and for a skewed sample it is a weak one.
  • This audits precision only. Four of this programme's corrections were to the analysis rather than the sample, and no interval-based scan would have caught any of them.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
overfitting
When a model learns the training data specifically rather than the pattern behind it, and so does well in training and badly on anything new.
rank
How many independent directions a set of numbers really uses. A low-rank structure is one that looks high-dimensional but is actually simple underneath.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
width
How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.