Research question

Is the task we are studying actually hard?

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means. More about this programme.

Before asking how a model solves something, it is worth asking whether the something is difficult at all -- and what the cheapest thing that also works would score.

Where this stands

Often it is not. A rule from 1990 with no parameters beats the trained model on the task most of these results were measured on, and what an intervention costs is set by the task's own structure.

This is the check that reframes everything else, and almost nobody runs it. If you are choosing a benchmark, run the dullest possible method on it first -- it takes minutes.

Still open · 4 published results bear on this question.

What this does not settle yet

Stated plainly, because the gaps are as much a part of the record as the answers:

  • Is there any task in the battery a GRU beats every classical baseline AND a transformer on?
  • Does the cost-class ordering survive at other lags? N12 gives reason to doubt it.

How these results fit together

Four records, and the reason this question exists at all is uncomfortable: a rule with no learned parameters, of a kind known since about 1990, solves the main task here perfectly, while the network spends substantial compute rediscovering it. That does not invalidate the other results -- they honestly describe how these networks find that mechanism -- but it sets a limit on what they can be about. So we built a battery of harder tasks with deliberate room above the ceiling, because a task the model already scores 99% on cannot show you that something made it slightly worse. Two records here are more interesting than the framing suggests. The cost of denying a model part of its task turns out to be set by *which* part of the task's repeating structure is removed, and those cost groupings move when the structure moves -- which means the task is telling us something about the model rather than the reverse. And the parts learned earliest are the most expensive to disrupt, which is where the idea of a prerequisite structure inside learning came from.

The results

Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.

A Rule From 1990

A rule with no parameters, invented before neural language models existed, solves our main task outright. That does not make our results wrong; it changes what they are about.

Read the record

A Task Battery With Designed Headroom

Two of four candidate tasks accepted. The rejections are the informative part: one is all-or-nothing per seed, the other is trivial for a recurrent model.

Read the record

Predicting The Cost From The Task

The first thing we have found here that can be predicted from the exercise alone, without looking at the model at all.

Read the record

The First Piece Is The Load-Bearing One

The part a model learns first costs the most to remove, and the part it struggles longest with costs the least.

Read the record

Back to all research questions

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.