Research record

Matched, and the Sum Still Halves

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: Is the task we are studying actually hard? – Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.

In plain English

What we asked. Our easy-start recipe sped up adding two remembered symbols but slowed learning a random lookup table. But those experiments used different numbers of symbols and different-sized models, so we matched them: the sum with exactly the table's 20 symbols and the same model.

What we found. The sum was still learned in about half the steps with the easy start. Ordinary training found it harder than one of the random tables and easier than the other, and the easy start slowed both of those.

Why it matters. So it is not how hard the task is, or how big the model is. It is whether the easier versions share a rule with the hard one.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 96 training runs of 12000 steps on sixteen fresh seeds. The design, a reframing made before any pilot, a disclosed calibration and the kill test were committed (98146fe, 95d107f) before any run.

Program v2 Bucket A, item A73. Decisive computation: analysis/sum_twenty.py. Output: analysis/sum_twenty.json.

The question

On a random 20-symbol lookup table the easy start was slower than plain training at every rate (A66, A70), and on a harder 24-symbol table too (A68). The sums it sped up were at 32 symbols, with a larger model input and output. With the symbol count and model matched to A66, does the sum still get the speed-up? (The item first proposed a sum behind randomly relabelled outputs; that was reframed before any pilot, because with learned embeddings and a linear readout a relabelled sum is the sum to this network.)

Same size, same model: the sum is sped up, random tables are slowed
Same size, same model: the sum is sped up, random tables are slowed. Three tasks with the same number of symbols and the same model: two random lookup tables and the sum. Below zero means the easy start was faster. The sum sits between the two tables in difficulty, yet it gets about half the steps while both tables are slowed. What decides it is whether the easy versions share a rule with the hard one.

Design: A66's loop with 20 symbols and target (x[t-2] + x[t-7]) mod 20; easy sums on (2,3)/(2,4)/(2,5) for 2000 steps; 12000 steps; sixteen fresh seeds; both arms at 0.003, 0.005, 0.008; best of each arm, paired. Kill test, fixed before execution: best mixed minus best plain includes zero or lies above it. Anchor, in code: A70's plain 0.008 run reproduces -- held.

Results

RateEasy start first: mean steps (median), solvedPlain: mean steps (median), solved
0.0033407 (3415), 166471 (5895), 15
0.0053171 (2960), 168392 (9295), 8
0.0083452 (2775), 1511558 (12000), 1
  • Best mixed (0.005) minus best plain (0.003): -3300.0 [-4617.4, -1982.6] steps; ratio of mean steps 0.490.

The kill test does not fire. With symbols and model matched to the random table, the easy start roughly halves the steps on the sum, as it did at 32 symbols (0.48-0.58).

Plain's best rate is the bottom of its grid. The script's edge check looked only at the top (and printed False); plain gets steadily worse as the rate rises, so a lower rate could help it. A75 runs plain at 0.002 and 0.0015.

What it says

Set beside the tables, this isolates structure from difficulty and from model size:

Task (20-24 symbols, same loop)Plain at its best (mean steps)Easy start minus plain
random table, 20 symbols (A70)4707+1976 (slower)
sum mod 20 (A73)6471-3300 (about half)
random table, 24 symbols (A68)9858+2997 (slower)

The sum sits between the two tables in difficulty for plain training, and gets the speed-up while both tables are slowed. So it is the shared rule, not the difficulty, the symbol count or the model size, that decides whether the easy start helps -- with the rate caveat on plain here (A75).

A calibration misled, and is disclosed. Two calibration seeds put plain's sum at about 3800 steps (slightly easier than the table); over sixteen seeds it is 6471 (harder). The conclusion does not depend on it -- the tables bracket the sum on both sides -- but the pilot's stated reason for calling the comparison clean ("slightly easier") was wrong.

What stands

  • A73: kill test does not fire. Sum mod 20: best mixed minus best plain -3300.0 [-4617.4, -1982.6] steps (ratio 0.490), against slow-downs on random tables either side of it in difficulty. Plain's best rate is its grid bottom (A75).

Limits

  • Plain's optimum not found (A75). One modulus; width 48; sixteen seeds.
  • Plain at 0.003 leaves one run unsolved at the 12000-step budget, counted at 12000.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

calibration
Working out an instrument's settings from runs whose answer you already know, so it can be used on a run whose answer you do not.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
width
How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.