The short answer: METR's 50% time horizon is the length of task, measured in human working time, that a frontier model completes successfully half the time. In March 2025 it sat around an hour, and it has roughly doubled every seven months since 2019. Read as a ceiling on unsupervised work, it explains why the exposure ratings light up on drafting and research and go quiet on running a case, a client, or a quarter.
What METR measures
METR, an evaluation lab, builds tasks with known human completion times, from a few seconds to many hours, and runs models against them. The 50% time horizon is the task length at which a model's success rate crosses fifty percent. Shorter tasks it usually completes; longer ones it usually does not.
In their March 2025 paper the frontier sat around an hour of human work. The measure has been doubling roughly every seven months since 2019. Those are their figures, from their task set, and the paper is careful about the limits of generalising from it. Even so, it is the clearest single line I know of through the question of what current AI can do on its own.
Read it as a ceiling on unsupervised work
A task that takes you four minutes sits well inside the horizon. Reformat this, summarise that, draft a first version of a routine document. The model will usually finish it, and a person can check it in less time than it took to write.
A task that takes three days, with judgement calls at hour six that change what you do at hour thirty, sits outside it. The individual steps are no harder. What changes is that errors compound, and nobody caught them in the middle. By hour thirty the work is built on a decision from hour six that nobody reviewed.
That is why the exposure ratings light up on drafting and research and go quiet on running a case, a client, or a quarter. The long tasks are not harder step by step. They are longer, and length is where unsupervised work breaks.
How it lines up with the exposure ratings
The Science exposure rubric asks whether a model could halve the time on a task without losing quality. Tasks that clear that bar are overwhelmingly the short, bounded ones: a document from a brief, a summary from a transcript, a memo from a statute. Tasks that fail it are the long, branching ones: manage this matter through to settlement, run this account for a year, deliver this project. The time horizon is a different measurement from a different lab, and it points at the same line.
That agreement is worth more than either number alone. When two independent methods sort tasks the same way, the sorting is probably real.
What it tells you to watch
The horizon moves. If it keeps doubling at the observed rate, the tasks inside it in two years are the ones that take you a full working day today. The tasks inside it in four years are the ones that take a week.
That gives you a concrete question about your own work: what is the longest task you do that nobody checks in the middle? If that task takes an hour, it is already inside the horizon. If it takes a day, it is roughly two years out on the current trend. If it takes a week, longer. None of that is a forecast about your job. It is a way of ordering your tasks by when the technology is likely to reach them, which is more useful than a ranking of your title.
The limits
METR's tasks are drawn mostly from software and research work, and the paper says the horizon may differ in other domains. Real jobs also have interruptions, ambiguous specifications, and stakeholders, none of which a benchmark task carries. Treat the horizon as a direction and a rough scale, not as a schedule. On those terms it is the best single number I have found for this question.
Frequently asked questions
What is METR's 50% time horizon?
The length of task, measured in how long a human takes, that a frontier AI model completes successfully half the time. METR reported it at roughly one hour in March 2025, doubling about every seven months since 2019.
Why do long tasks resist AI even when each step is easy?
Errors compound. A three-day task contains judgement calls early on that shape everything after. Without someone checking in the middle, a small early mistake propagates through the rest of the work.
How should I use the time horizon for my own job?
List your tasks by how long they take without anyone reviewing them. Tasks under an hour are already inside the horizon. On the current trend, day-long tasks are roughly two years out. That ordering tells you where change is likely to arrive first.
Get your free AI Job Risk Score. Tell us your job title and how you actually spend your time, and we will show you which of your tasks are exposed today. Free. 60 seconds.
Get my AI risk now