Research record

A Text Curriculum on a Small Model

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Part of a bigger question: What actually makes training cheaper? – One thing has worked: stopping part of the training early saved about 7% with no loss of quality. Everything else tested has been matched by a simpler or cheaper method -- and in two cases the clever method was only winning because it was quietly being given more.

In plain English

What we asked. A 2026 paper found that training a language model on its easiest documents first saved up to 45% of the steps, but it never compared that with simply warming up the learning rate. We tried both on a small model reading this project's own writing.

What we found. Neither helped. A plain, constant learning rate was ahead of both the whole way. But when we looked at which documents counted as easiest, they were mostly two near-identical automatically generated lists, so the easy-first recipe spent its start memorising boilerplate.

Why it matters. So this run says little about the paper. We are rerunning it with a properly varied set of easy documents. The lesson worth keeping: check what your difficulty measure actually picks before trusting it.

The rest of this page is the technical record: the design, every number, and the limits. It is written for a reviewer, and you do not need it to have understood the result above.

New here? How to read a research record
  • Start at the verdict. Every record states, before the experiment was run, what result would have made us abandon the idea. That is the "kill test". Then it says whether the test fired. Nothing gets reinterpreted after the fact.
  • Numbers in square brackets are uncertainty. 23.4 [18.1, 28.7] means our best estimate is 23.4 and the true value is probably somewhere in that range. If a range includes zero, we cannot claim an effect.
  • Results that rule an idea out are kept. Roughly half of what is published here says an approach did not work, including plenty of our own. Those pages are the output, not a shortfall: knowing which direction is a dead end is what lets the next experiment go somewhere better, and most of what we now rely on came out of understanding why something else fell short. Work that only publishes what worked is not measuring anything.
  • Read the Limits section. Every record ends with what it does not show. It is the most honest part of any experiment and usually the shortest.
  • Pro tip: the figures near the top are designed to carry the result on their own. If you read nothing else, read the caption under each one, which says what it shows and what to take from it.
EXPLORATORY. Not a preregistered study. Local CPU, 48 training runs of 4000 steps on sixteen seeds. The design, a disclosed calibration and the kill test were committed (8c276e6) before any run. A post-hoc analysis, labelled as such, was added after the endpoint was read.

Program v2 Bucket A, item A72. Decisive computation: analysis/text_curriculum.py; post-hoc: analysis/text_curriculum_posthoc.py. Outputs: analysis/text_curriculum.json, analysis/text_curriculum_posthoc.json.

The question

arXiv 2506.11300 (EACL 2026) reports that ordering pretraining data easy-first as a warm-up reaches a random baseline's best loss in 18-45% fewer steps, with every arm at a fixed learning rate and no learning-rate warm-up. On our copy task (A33) that comparison credited a curriculum with what a rate warm-up does. On text, does a curriculum beat a plain learning-rate warm-up of the same length?

On a small text model, a constant learning rate was best from start to finish
On a small text model, a constant learning rate was best from start to finish. Three ways to start training a small model on text: a constant learning rate, the easiest documents first, and a learning rate that ramps up. The easiest documents turned out to be mostly two near-identical generated indexes. Neither starting trick beat the plain constant rate here. The curriculum's easy pool was a poor one, so this does not test the published result it was aimed at; a rebuilt version is queued.

Design: a one-layer GRU, width 128, character-level, on this repository's prose at a pinned revision (48897 chunks of 65 characters, 96 symbols, hash-checked); difficulty is each document's gzip compression ratio. Three arms at peak 0.003, 4000 steps, sixteen seeds: constant (random chunks, constant rate -- the paper's baseline); curriculum (the easiest third by document ratio for the first 1000 steps, then random -- the paper's method); rate warm-up (random chunks, rate rising linearly over 1000 steps -- our control). Endpoint: steps to reach the constant arm's final held-out loss. Kill test, fixed before execution: curriculum minus rate warm-up includes zero or lies above it. Anchor, in code: the corpus hash -- held.

Results

ArmMean steps to the constant arm's final lossReached it (of 16)Final held-out loss
constant3869121.3199
curriculum399411.3272
rate warm-up391951.3242
  • Curriculum minus rate warm-up: +75.0 [-9.2, +159.2] steps -- inside one evaluation interval (100 steps).
  • Curriculum minus constant: +125.0 [+50.4, +199.6]; rate warm-up minus constant: +50.0 [-70.7, +170.7].

The kill test fires. The curriculum did not beat the rate warm-up.

Post-hoc, labelled as such (paired held-out loss differences, from the committed output):

At step 1000At step 2000Mean over the runFinal
curriculum - constant+0.155+0.013+0.046+0.007
rate warm-up - constant+0.073+0.016+0.108+0.004
curriculum - rate warm-up+0.082-0.002-0.062+0.003

The constant-rate baseline is ahead of both at every point. The curriculum is far behind while it trains on the easy pool (its held-out loss is 1.564 at step 1000 against 1.409) and nearly catches up afterwards; the warm-up starts slower still and also nearly catches up.

What it says, and what it cannot

The paper's direction did not appear at this scale: the curriculum was slower than the constant baseline, not faster. So this run cannot say whether the paper's gain is a warm-up's -- there was no gain to attribute. The preregistered kill test fired, and it should be read as "no curriculum benefit here", not as evidence about the paper.

The design was weaker than intended, and the post-hoc check shows why. The "easiest third" by document compression ratio was four documents, and 46% of it was two near-identical auto-generated figure indexes (the light and dark docs/figures/README.md), with the rest the programme file and the decision log. Training on that pool for 1000 steps is not the varied easy-first curriculum the paper describes; a run that first memorises boilerplate and then moves to the real distribution should be expected to lag. That is a design fault of this pilot, found after the run, and it is the main reason not to read anything into the paper from it. A74 repeats the test with the curriculum's easy pool built without generated files and with no single document above a small share.

Separately, a learning-rate warm-up did not help this model either -- at this rate and length, the constant rate was simply best.

What stands

  • A72: kill test fires (curriculum minus warm-up +75.0 [-9.2, +159.2] steps, within one evaluation interval). The paper's curriculum gain did not reproduce at this scale; the easy pool was four documents, half of them duplicated generated indexes, so the run is weak evidence about curricula and none about the paper.

Limits

  • A one-layer GRU of width 128 on 3 MB of text for 4000 steps -- far from the paper's scale. One peak rate.
  • The endpoint (steps to the constant arm's own final loss) sits at the end of the run by construction; the post-hoc loss table is the more informative view and is labelled as post-hoc.
  • The easy pool's composition, above.

Terms on this page

Every piece of vocabulary this record uses, in plain language. Generated from the text above, so it cannot drift out of step with it.

baseline
The thing you compare against. A result without one is not a result.
calibration
Working out an instrument's settings from runs whose answer you already know, so it can be used on a run whose answer you do not.
curriculum
Training on easier examples first and harder ones later, like a school syllabus, rather than on everything at once.
evaluation interval
The gap between successive checks of a model during training. It is the smallest difference in timing a measurement can resolve, so two results closer together than one interval cannot be told apart.
GRU
Gated Recurrent Unit. A compact design for processing sequences one item at a time, with internal switches controlling what it keeps in memory.
held-out
Data the model was never trained on, kept back specifically to test it. Scoring a model on data it has already seen measures memorisation, not learning.
kill test
A condition written down before running the experiment that says what result would make us abandon the idea. Fixing it in advance is what stops a disappointing result being reinterpreted as an encouraging one.
learning rate
How big a step training takes each time it updates the model. Too small and nothing happens; too big and it never settles.
post hoc
Worked out after the fact, rather than decided in advance. We report such checks separately and never let them decide a result, because it is far too easy to find a pattern once you already know the answer.
seed
The number that fixes all the randomness in a training run. Same seed, same run. Running several seeds is how you tell a real effect from a lucky one.
width
How many internal numbers a model uses at each layer. The usual way we vary model size in these experiments.

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.