A Text Curriculum on a Small Model
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Part of a bigger question: What actually makes training cheaper? – One thing has worked: stopping part of the training early saved about 7% with no loss of quality. Everything else tested has been matched by a simpler or cheaper method -- and in two cases the clever method was only winning because it was quietly being given more.
In plain English
What we asked. A 2026 paper found that training a language model on its easiest documents first saved up to 45% of the steps, but it never compared that with simply warming up the learning rate. We tried both on a small model reading this project's own writing.
What we found. Neither helped. A plain, constant learning rate was ahead of both the whole way. But when we looked at which documents counted as easiest, they were mostly two near-identical automatically generated lists, so the easy-first recipe spent its start memorising boilerplate.
Why it matters. So this run says little about the paper. We are rerunning it with a properly varied set of easy documents. The lesson worth keeping: check what your difficulty measure actually picks before trusting it.
The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.
New here? How to read a research record
- Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
- Numbers in square brackets are uncertainty.
23.4 [18.1, 28.7]means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect. - Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
- Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
- Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU,48training runs of4000steps on sixteen seeds. The design, a disclosed calibration and the kill test were committed (8c276e6) before any run. A post-hoc analysis, labelled as such, was added after the endpoint was read.
Program v2 Bucket A, item A72. Decisive computation: ; post-hoc: analysis/text_curriculum.py. Outputs: analysis/text_curriculum_posthoc.py, analysis/text_curriculum.json.analysis/text_curriculum_posthoc.json
The question
arXiv 2506.11300 (EACL 2026) reports that ordering pretraining data easy-first as a warm-up reaches a random baseline's best loss in 18-45% fewer steps, with every arm at a fixed learning rate and no learning-rate warm-up. On our copy task (A33) that comparison credited a curriculum with what a rate warm-up does. On text, does a curriculum beat a plain learning-rate warm-up of the same length?
Design: a one-layer GRU, width 128, character-level, on this repository's prose at a pinned revision (48897 chunks of 65 characters, 96 symbols, hash-checked); difficulty is each document's gzip compression ratio. Three arms at peak 0.003, 4000 steps, sixteen seeds: constant (random chunks, constant rate -- the paper's baseline); curriculum (the easiest third by document ratio for the first 1000 steps, then random -- the paper's method); rate warm-up (random chunks, rate rising linearly over 1000 steps -- our control). Endpoint: steps to reach the constant arm's final held-out loss. Kill test, fixed before execution: curriculum minus rate warm-up includes zero or lies above it. Anchor, in code: the corpus hash -- held.
Results
| Arm | Mean steps to the constant arm's final loss | Reached it (of 16) | Final held-out loss |
|---|---|---|---|
| constant | 3869 | 12 | 1.3199 |
| curriculum | 3994 | 1 | 1.3272 |
| rate warm-up | 3919 | 5 | 1.3242 |
- Curriculum minus rate warm-up:
+75.0[-9.2, +159.2]steps -- inside one evaluation interval (100steps). - Curriculum minus constant:
+125.0[+50.4, +199.6]; rate warm-up minus constant:+50.0[-70.7, +170.7].
The kill test fires. The curriculum did not beat the rate warm-up.
Post-hoc, labelled as such (paired held-out loss differences, from the committed output):
At step 1000 | At step 2000 | Mean over the run | Final | |
|---|---|---|---|---|
| curriculum - constant | +0.155 | +0.013 | +0.046 | +0.007 |
| rate warm-up - constant | +0.073 | +0.016 | +0.108 | +0.004 |
| curriculum - rate warm-up | +0.082 | -0.002 | -0.062 | +0.003 |
The constant-rate baseline is ahead of both at every point. The curriculum is far behind while it trains on the easy pool (its held-out loss is 1.564 at step 1000 against 1.409) and nearly catches up afterwards; the warm-up starts slower still and also nearly catches up.
What it says, and what it cannot
The paper's direction did not appear at this scale: the curriculum was slower than the constant baseline, not faster. So this run cannot say whether the paper's gain is a warm-up's -- there was no gain to attribute. The preregistered kill test fired, and it should be read as "no curriculum benefit here", not as evidence about the paper.
The design was weaker than intended, and the post-hoc check shows why. The "easiest third" by document compression ratio was four documents, and 46% of it was two near-identical auto-generated figure indexes (the light and dark docs/figures/README.md), with the rest the programme file and the decision log. Training on that pool for 1000 steps is not the varied easy-first curriculum the paper describes; a run that first memorises boilerplate and then moves to the real distribution should be expected to lag. That is a design fault of this pilot, found after the run, and it is the main reason not to read anything into the paper from it. A74 repeats the test with the curriculum's easy pool built without generated files and with no single document above a small share.
Separately, a learning-rate warm-up did not help this model either -- at this rate and length, the constant rate was simply best.
What stands
- A72: kill test fires (curriculum minus warm-up
+75.0[-9.2, +159.2]steps, within one evaluation interval). The paper's curriculum gain did not reproduce at this scale; the easy pool was four documents, half of them duplicated generated indexes, so the run is weak evidence about curricula and none about the paper.
Limits
- A one-layer GRU of width
128on3MB of text for4000steps -- far from the paper's scale. One peak rate. - The endpoint (steps to the constant arm's own final loss) sits at the end of the run by construction; the post-hoc loss table is the more informative view and is labelled as post-hoc.
- The easy pool's composition, above.
Terms on this page
Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.
- baseline
- The thing you compare against. A result without one is not a result.
- calibration
- Working out an instrument's settings from runs whose answer you already know, so it can be used on a run whose answer you do not.
- curriculum
- Training on easier examples first and harder ones later, like a school syllabus, rather than on everything at once.
- evaluation interval
- The gap between successive checks of a model during training. It is the smallest difference in timing a measurement can resolve, so two results closer together than one interval cannot be told apart.
- GRU
- Gated Recurrent Unit. A compact design for processing sequences one item at a time, with internal switches controlling what it keeps in memory.
- held-out
- Data the model was never trained on, kept back specifically to test it. Scoring a model on data it has already seen measures memorisation, not learning.
- kill test
- A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
- learning rate
- How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
- post hoc
- Worked out after the fact, rather than decided in advance. We report such checks separately and never let them decide a result, because it is far too easy to find a pattern once you already know the answer.
- seed
- The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
- width
- How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.