BriefPulse Practical AI · Working notes on AI you can actually use. RSS · BriefPulse network
BriefPulse Practical AI

What changed in AI, what it is useful for, and what you can do with it.

17 September 2026

Brief

METR proposes an 'expenditure horizon' to measure an AI agent's optimization ability

METR proposes a measure of an AI agent's optimization ability, illustrated with the NanoGPT speedrun. The source is a short method note rather than a tool or a result, so there is no number to apply yet.

The METR post proposes a measure of an AI agent's optimization ability called an 'expenditure horizon.' The problem it addresses is cost accounting: comparing human and agent performance on the same optimization task means weighing token cost, experiment compute cost and human labor cost, which the post names as one difficulty in measuring AI's ability to accelerate AI R&D.

The proposed measure is the crossing point of two curves — performance as a function of cost for humans and for agents — giving the budget at which humans become more cost-effective than AIs. The method is illustrated with data from the NanoGPT speedrun.

Our reading

For an evaluation desk, the interesting move here is bundling spend into the capability metric instead of reporting accuracy or speed alone — which is closer to how anyone actually deciding whether to delegate work has to think. It matters to teams running agents on optimization or engineering tasks, and to anyone building an evaluation harness, because cost curves are a different measurement dis…

What to do or watch

Watch for the full method: the curves, how each cost category is priced, and whether the crossing point holds up outside the NanoGPT speedrun. Until then the unresolved question is at what budget, on a task you actually run, handing the work to a human becomes the cheaper option.

Source details and supporting facts

Each line is stated by the page named above it.

Stated by metr.org

  • The post proposes a measure of an AI agent's optimization ability with an 'expenditure horizon.'
  • One difficulty in measuring AI's ability to accelerate AI R&D is accounting for token cost, experiment compute cost and human labor cost.
  • The method is illustrated with data from the NanoGPT speedrun.

Sources

  1. METRText stored 17 September 2026

How this story was checked. Written from the 1 page listed above, stored 17 September 2026; claims checked against that stored text on 17 September 2026.

What that means
  • 3 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
  • Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
  • The check reads stored text only: no claim rests on a fresh look that did not happen.
  • Where the reporting was silent, the text says so instead of filling the gap.

More from Practical AI