CourionAI
EN
Newsletter
← All news
research 2 min read

METR has a new way to ask whether an AI agent is actually cheaper than a person

The evaluation lab proposes the expenditure horizon: the budget at which a human becomes more cost-effective than an AI agent on the same optimization task. In a first test, agents crossed over somewhere between zero and $3,000.

Risograph illustration of two curved lines crossing on a faint graph grid above a row of stacked coin discs

Most AI benchmarks answer a yes-or-no question: did the model solve the task. METR, the nonprofit lab that evaluates frontier models for the big labs, has proposed a metric that answers a more useful one: at what budget does a human become the better buy. They call it the expenditure horizon, and it is defined as the dollar amount at which the improvement an agent delivers equals the improvement a human would deliver for the same money.

The test case is the NanoGPT speedrun, a long-running community competition to train a small language model as fast as possible. It is a good testbed because humans have already optimized it hard, so there is no easy headroom left. METR first measured what human progress costs: on the margin, squeezing out each additional one percent improvement takes roughly $2,500 of human labor. Then they ran agents at the same problem and plotted score against spend. After more than $10,000 of agent expenditure, they estimate the expenditure horizon sits somewhere between $0 and $3,000. In plain terms: on this particular problem, agents are competitive with humans for small budgets, and humans pull ahead once you are spending serious money.

What makes this worth attention is less the result than the shape of the measurement. Existing AI research benchmarks tend to report a pass or fail against a human who worked for eight or forty hours, which throws away most of the information and needs many runs to detect a difference between two models. A continuous curve of score against dollars is statistically sharper, it folds in the cost of the compute the agent burns while experimenting, and it can be pointed at problems humans have already picked clean rather than textbook exercises. METR is careful about what it does not show: this is one problem, the agent runs are preliminary, and the human cost estimate is rough. Nobody should read a single expenditure horizon as a verdict on AI research capability.

What this means for you: If you are curious about AI but tired of benchmark scores that mean nothing to you, this is a metric you can actually reason about, because it is denominated in money rather than percentage points. If you already use AI agents for real work, the practical takeaway is the one METR’s curve implies: agents tend to be excellent value for the first stretch of a task and worse value as the problem gets deep and expensive. Budgeting for that, rather than assuming an agent is uniformly cheaper, is the useful adjustment. Expect labs to start reporting scaling curves rather than single scores, because once someone measures cost this cleanly, it becomes hard to go back.

Sources

Source: https://metr.org/blog/2026-07-21-expenditure-horizon/

Next story

Nvidia may guarantee $250 billion so OpenAI can rent a data center

The Wall Street Journal reports Nvidia is in talks to backstop about $250 billion of financing for a 10-gigawatt campus in Ohio that OpenAI would lease. A chip vendor underwriting its customer's rent is unusual, and it says something about who carries the risk in this boom.

Risograph illustration of a vast low server hall propped up from below by an oversized lattice of steel scaffolding, cooling towers in the distance