A Clock, Not a Warning
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Part of a bigger question: What actually happens at the moment a model learns? – It builds machinery rather than selecting it, working through the task in a reproducible order and trying a simpler wrong rule on the way. The visible training curve cannot tell you which is happening.
In plain English
What we asked. Last time the model's internal dynamics moved towards the edge of chaos at the same early moment in every run. To see whether that move warns of learning, we changed the learning speed so learning happened at very different times.
What we found. The move shifted too, but always stayed about 30 percent of the way to learning, and was identical between runs at the same speed. It tracks how far training has got, like a clock, and tells you nothing about when a particular run will learn.
Why it matters. The lesson: before calling something an early warning, move the event and check the warning moves with it, and varies with it run to run.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU,24training runs with Lyapunov measurements. The design and kill test were committed (f992f21) before any run.
Program v2 Bucket N, item N29. Decisive computation: . Output: analysis/edge_of_chaos_moved.py.analysis/edge_of_chaos_moved.json
The question
N28 found the GRU's recurrence covers half its move towards the edge of chaos by step 50 on every receiver, about 100 steps before the transition, and that its preregistered test passed only because a constant early step beats a random one. A lead means something only if it moves with the event (K1). N29 moves the event by running rates 0.001, 0.002 and 0.004, and measures the shift ratio: the mean half-rise shift between the slowest and fastest rate over the mean transition shift (a ratio of means, bootstrap interval).
Kill test, fixed before execution: the shift ratio's interval lies entirely below 0.5. Anchor: at 0.002 every receiver reproduces J8's control. It holds.
Result: the kill test fires
| Rate | Half-rise (eight receivers) | Transition (eight receivers) | Mean half-rise / mean transition |
|---|---|---|---|
0.001 | 85-90 | 254-290 (mean 268.4) | 0.32 |
0.002 | 45-50 | 148-167 (mean 154.2) | 0.30 |
0.004 | 25-30 | 91-104 (mean 95.8) | 0.28 |
Shift ratio 0.348 [0.334, 0.362]: the half-rise moves about a third as far as the transition.
What it is instead
The half-rise is not fixed -- it moves with the learning rate -- but it sits at a nearly constant fraction of the transition time (0.28-0.32), because both scale with the inverse of the rate: the optimiser covers the same ground in fewer steps. And within a rate it carries no information about the individual run: across eight receivers the half-rise varies by one measurement step (5 steps) while the transition varies by 15-36. The approach to the edge of chaos is a clock of the optimiser's progress, reached a third of the way to learning whatever the run, not a precursor that sees a particular run's transition coming.
What stands
- Kill test fires: the half-rise follows only
0.35of the transition's move. - The recurrence approaches the edge of chaos on a fixed schedule in optimiser time -- about
30%of the way to the transition at every rate -- and says nothing about which run will learn sooner. - With N28: the reverse of arXiv
2609.19288's direction, early, and uninformative about timing. The dynamical thread closes here for this substrate.
Limits
- One task, one width; the grid is
5steps, so finer per-run variation is not resolved.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- bootstrap
- A way of estimating how uncertain a number is by repeatedly resampling the data you already have. Useful when the usual formulas do not apply.
- GRU
- Gated Recurrent Unit. A compact design for processing sequences one item at a time, with internal switches controlling what it keeps in memory.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- learning rate
- How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.