Can AI Agents Complete Multi-Step Tasks Unsupervised?
The short answer
Yes for short, well-defined tasks with automatic checks; not yet for long, ambiguous or irreversible ones. The best public measurement, METR’s task-completion time horizons, shows frontier agents handling software tasks worth several hours of expert work about half the time, but reaching 80% reliability only around the one-to-three-hour mark. For anything where a 1-in-5 failure is unacceptable, the work needs a human checkpoint.
What the data says
METR estimates how long a skilled human takes on each of 228 software, ML and security tasks (Time Horizon 1.1, released January 29, 2026), then fits each agent’s success rate against task length.
| Model (METR TH1.1) | 50% time horizon | 80% time horizon |
|---|---|---|
| Claude Mythos Preview (early) | ~17 hours* | ~3.1 hours |
| Claude Opus 4.6 | ~12 hours | ~1.2 hours |
| Gemini 3.1 Pro | ~6.4 hours | ~1.5 hours |
| GPT-5.4 | ~5.7 hours | ~54 minutes |
| GPT-5.3-Codex | ~5.8 hours | ~55 minutes |
*METR notes measurements above 16 hours are unreliable with its current task suite. Source: METR’s published results file, last page update May 8, 2026. Models released since (Claude Opus 5.5, GPT-6, Gemini 4 Argon) have no METR measurement yet.
Three things stand out:
- The 80% horizon is 4–10× shorter than the 50% horizon. “Can sometimes do a day’s task” and “reliably does an hour’s task” are both true of the same model.
- Progress is fast. METR’s fit puts the doubling time of the 50% horizon at about 129 days for models since 2023 (confidence interval 104–158 days).
- The tasks are easier than real work. METR says its tasks are self-contained and well-specified, closer to what “a new hire or a remote contractor” could do, and that agent performance “drops substantially” when results are scored holistically rather than by automated tests.
Agents versus scripts
If a task is fixed and repeatable — rotate logs, sync two databases, run the nightly build — a deterministic script or workflow (cron, n8n, Make, a CI job) beats an agent on cost, predictability and auditability. Agents earn their place where steps can’t be specified in advance: triaging an unfamiliar bug, researching across sources, adapting to a page that changed. A common 2026 design puts the agent where judgement is needed and keeps everything else deterministic.
Where unsupervised agents work today
- Coding tasks with tests: fix a failing test, add an endpoint with its tests, upgrade a dependency and make CI pass. The test suite is the supervisor.
- Research and summarisation where a person reads the result before acting.
- Data extraction into a schema with validation and confidence thresholds.
- Background coding agents (Codex cloud, Cursor cloud agents, Claude Code) that open a pull request for review rather than merging.
Where they still need a human
- Irreversible actions: payments, deletions, production deploys, sending email to customers.
- Ambiguous goals: “improve onboarding” without a measurable target.
- Long chains without checkpoints: a 50-step workflow at 98% per-step success finishes cleanly only about a third of the time (0.98⁵⁰ ≈ 0.36).
- Untrusted inputs: web pages and documents can carry prompt injections that redirect an agent’s tools.
How to deploy agents so they’re reliable enough
- Define success in code: tests, JSON schemas, assertions. If you can’t check it automatically, plan for review.
- Keep tasks short: split long jobs into steps an agent completes at high reliability, with saved state between them.
- Least privilege: scoped API keys, read-only by default, a sandbox for code execution.
- Approval gates on irreversible or expensive actions.
- Observe everything: log each tool call and decision so failures can be diagnosed.
- Measure your own success rate on 50–100 real tasks before reducing oversight. Benchmarks tell you the trend; only your tasks tell you your reliability.
Related: best AI agent control planes, best tools for running multiple AI coding agents, AWS Bedrock AgentCore alternatives.
Last verified: October 6, 2026. Time horizons converted from METR’s published minutes and rounded.