What is an AI agent, and can it finish the job yet?

Everyone sells agents; few define them. A plain-language guide to what an AI agent actually is, and what the best independent evidence says about how long one can work before it breaks.

By Yash Malviya

Published

Hands typing on a laptop with a blank screen on a white desk, smartphone nearby
Photo: MART PRODUCTION / Pexels

Start by refusing the demo's definition

An AI agent is a model given a goal, a set of tools and permission to loop: it plans, acts through software such as browsers, terminals and APIs, checks its own results and tries again, without a human prompting every step. That is the whole definition. A chatbot answers you. An agent acts for you. If the product needs you to click approve at each step, it is a workflow with good branding, and if it collapses the first time reality deviates from the script, it is a demo.

NVIDIA's engineers, writing about agent training data on Hugging Face in July, put the bar plainly: "An agent that can't recover from a broken API call, or a workflow it has never seen, is not really an agent."

The only independent yardstick worth your time

The marketing question is whether agents are magic. The measurable question is how long one can work on its own before it breaks, and the best independent answer comes from METR, a non-profit evaluator whose "time horizon" metric estimates the length of task, measured in expert-human working time, that a model completes at a given reliability.

METR's Time Horizon 1.1 update, published on 29 January 2026 across a suite of 228 tasks, put the longest measured horizon at Claude Opus 4.5: about 320 minutes at 50% reliability, with a wide error range of 170 to 729 minutes. GPT-5 measured around 214 minutes. More important than any single number is the curve: the horizon has doubled roughly every seven months across six years of models, and roughly every three months for models released since 2024. METR is candid about the limits, noting that "these confidence intervals are still very wide" and that the trend is "somewhat sensitive to task composition". The measured models also predate this autumn's GPT-6 and Opus 5.5 wave, so read the finding as a curve, not a snapshot.

“An agent that can't recover from a broken API call, or a workflow it has never seen, is not really an agent.”

NVIDIA engineers, 'Data for Agents', Hugging Face blog, 8 Jul 2026

Half the time is the operative phrase. A model that completes five-hour expert tasks at a coin-flip rate is a genuine research milestone and an operations nightmare in the same breath. Nobody staffs a role with someone who silently fails every second assignment; you staff around them with review, and that is exactly what sensible agent deployments do.

A person creates a flowchart diagram with red pen on a whiteboard, detailing plans and budgeting
The useful question is not what an agent is called, but how long one can work before it breaks. Photo: Christina Morillo / Pexels

What the vendors' own charts admit

Read the current launch posts closely and the honest numbers are hiding in plain sight. OpenAI's GPT-6 material shows its best configuration completing 33.2% of AutomationBench's end-to-end business workflows at $0.27 per task, and 56.4% on the Agents' Last Exam professional suite. Those are vendor-run evaluations chosen to flatter, and they still describe a system that does not finish most business workflows it starts. The pitch, stripped of styling, is that it fails cheaper than the competition.

Anthropic's launch material for Opus 5.5 leans on testimonials, including a developer at legal-software firm Clio reporting a run that "stayed on task for over 18 hours... required minimal reworking". Real, attributed, and vendor-selected: the best run makes the launch page, the median run does not.

The economics point the same direction. OpenAI disclosed that, valued at API prices, its median researcher burns more than $600 of tokens a day, with the 90th percentile past $7,000. Agents are token furnaces, which is why September's API price cuts are aimed squarely at them.

“It stayed on task for over 18 hours... required minimal reworking.”

Sean Heintz, staff software developer, Clio, in Anthropic's Opus 5.5 launch materials

Where agents earn their keep today

The pattern behind every agent success story is the same: the work was agent-shaped. Output that can be verified mechanically, by tests, compilers or a submitted form. Actions that can be undone. Rich tool access. Tolerance for retries. Coding leads the field not because models love code but because verification is built into the job, which is why every lab demos a coding agent first.

The failure pattern is equally consistent: irreversible, externally visible actions taken without a checkpoint, and tasks where checking the agent's work costs as much as doing it yourself. If you are evaluating a vendor, ask for the completion rate on your tasks rather than a benchmark, ask precisely what happens on failure, insist on human gates wherever an action cannot be undone, and put a quarterly re-test in the calendar, because the underlying curve moves in months, not years.

Our take

Agents are the rare AI story where the honest numbers are impressive without the marketing. An independent, methodologically transparent measure says autonomous working time doubles every few months, and the vendors' own charts, read plainly, admit today's systems abandon most long workflows they begin. Both facts are true and neither cancels the other. Deploy narrow, verify always, re-test quarterly, and ignore anyone selling a digital employee. What is on offer today is a fast, tireless assistant with a coin-flip memory for finishing, on the steepest improvement curve in software.

Frequently asked questions

What is an AI agent?

An AI agent is a model given a goal, a set of tools and permission to loop: it plans, acts through software such as browsers, terminals and APIs, checks its own results and tries again, without a human prompting every step. A chatbot answers you, while an agent acts for you. If a product needs you to click approve at every step, the article calls it a workflow with good branding.

How long can the best AI agents work on their own before they break?

METR's Time Horizon 1.1 update, published on 29 January 2026 across 228 tasks, put the longest measured horizon at Claude Opus 4.5: about 320 minutes at 50% reliability, meaning tasks that take human experts over five hours, completed about half the time, with a wide error range of 170 to 729 minutes. GPT-5 measured around 214 minutes. METR is candid that these confidence intervals are still very wide.

How fast is this time horizon improving?

According to METR, the horizon has doubled roughly every seven months across six years of models, and roughly every three months for models released since 2024. The article stresses the curve matters more than any single number. It also notes the measured models predate the autumn's GPT-6 and Opus 5.5 wave, so the finding should be read as a curve, not a snapshot.

Do the vendors' own numbers show agents finishing most tasks?

No. OpenAI's GPT-6 material shows its best configuration completing 33.2% of AutomationBench's end-to-end business workflows at $0.27 per task, and 56.4% on the Agents' Last Exam. These are vendor-run evaluations chosen to flatter, and they still describe a system that does not finish most business workflows it starts. The article says the pitch, stripped of styling, is that it fails cheaper than the competition.

Where do AI agents earn their keep today?

The pattern behind agent success is that the work is agent-shaped: output that can be verified mechanically by tests, compilers or a submitted form, actions that can be undone, rich tool access and tolerance for retries. Coding leads the field because verification is built into the job. The failure pattern is the opposite: irreversible, externally visible actions taken without a checkpoint, and tasks where checking the agent's work costs as much as doing it yourself.

Sources

What each one is, and whose it is.

  1. 1

    Time Horizon 1.1, METR (January 29, 2026)

    BenchmarkIndependent of the vendor
  2. 2

    Introducing GPT-6 Sol and Luna, OpenAI (September 22, 2026)

    Vendor announcement
  3. 3

    Introducing Claude Opus 5.5, Anthropic (September 22, 2026)

    Vendor announcement
  4. 4

    Data for Agents, NVIDIA on Hugging Face (July 8, 2026)

    Vendor announcement