Task horizon is the metric that maps to work

March 19, 2025
Loading the Elevenlabs Text to Speech AudioNative Player...

METR chart: length of tasks AI agents complete autonomously, doubling over time

METR's time-horizon paper asks a better question than "what is the SWE-bench score." How long a task, measured in expert-human time, can the agent complete with 50% reliability? Today's frontier agents are near-certain on work that takes a person a few minutes, and rare on work that takes several hours. Fit that into one number: a time horizon.

That horizon has been doubling roughly every seven months for six years. The 50% line is not "the agent can do your job." It is the length of task where a coin flip is the success rate. 80% is the more honest bar for anything you would leave unsupervised.

This is the measurement frame that matches production. We do not need another exam. We need to know whether the agent can carry a sequence of actions for as long as the job actually takes, and how often it still holds the thread at the end.