AI agents · OpenClaw · self-hosting · automation

Quick Answer

Can AI Agents Complete Multi-Step Tasks Unsupervised?

Published:

The short answer

Yes for short, well-defined tasks with automatic checks; not yet for long, ambiguous or irreversible ones. The best public measurement, METR’s task-completion time horizons, shows frontier agents handling software tasks worth several hours of expert work about half the time, but reaching 80% reliability only around the one-to-three-hour mark. For anything where a 1-in-5 failure is unacceptable, the work needs a human checkpoint.

What the data says

METR estimates how long a skilled human takes on each of 228 software, ML and security tasks (Time Horizon 1.1, released January 29, 2026), then fits each agent’s success rate against task length.

Model (METR TH1.1)50% time horizon80% time horizon
Claude Mythos Preview (early)~17 hours*~3.1 hours
Claude Opus 4.6~12 hours~1.2 hours
Gemini 3.1 Pro~6.4 hours~1.5 hours
GPT-5.4~5.7 hours~54 minutes
GPT-5.3-Codex~5.8 hours~55 minutes

*METR notes measurements above 16 hours are unreliable with its current task suite. Source: METR’s published results file, last page update May 8, 2026. Models released since (Claude Opus 5.5, GPT-6, Gemini 4 Argon) have no METR measurement yet.

Three things stand out:

  1. The 80% horizon is 4–10× shorter than the 50% horizon. “Can sometimes do a day’s task” and “reliably does an hour’s task” are both true of the same model.
  2. Progress is fast. METR’s fit puts the doubling time of the 50% horizon at about 129 days for models since 2023 (confidence interval 104–158 days).
  3. The tasks are easier than real work. METR says its tasks are self-contained and well-specified, closer to what “a new hire or a remote contractor” could do, and that agent performance “drops substantially” when results are scored holistically rather than by automated tests.

Agents versus scripts

If a task is fixed and repeatable — rotate logs, sync two databases, run the nightly build — a deterministic script or workflow (cron, n8n, Make, a CI job) beats an agent on cost, predictability and auditability. Agents earn their place where steps can’t be specified in advance: triaging an unfamiliar bug, researching across sources, adapting to a page that changed. A common 2026 design puts the agent where judgement is needed and keeps everything else deterministic.

Where unsupervised agents work today

  • Coding tasks with tests: fix a failing test, add an endpoint with its tests, upgrade a dependency and make CI pass. The test suite is the supervisor.
  • Research and summarisation where a person reads the result before acting.
  • Data extraction into a schema with validation and confidence thresholds.
  • Background coding agents (Codex cloud, Cursor cloud agents, Claude Code) that open a pull request for review rather than merging.

Where they still need a human

  • Irreversible actions: payments, deletions, production deploys, sending email to customers.
  • Ambiguous goals: “improve onboarding” without a measurable target.
  • Long chains without checkpoints: a 50-step workflow at 98% per-step success finishes cleanly only about a third of the time (0.98⁵⁰ ≈ 0.36).
  • Untrusted inputs: web pages and documents can carry prompt injections that redirect an agent’s tools.

How to deploy agents so they’re reliable enough

  1. Define success in code: tests, JSON schemas, assertions. If you can’t check it automatically, plan for review.
  2. Keep tasks short: split long jobs into steps an agent completes at high reliability, with saved state between them.
  3. Least privilege: scoped API keys, read-only by default, a sandbox for code execution.
  4. Approval gates on irreversible or expensive actions.
  5. Observe everything: log each tool call and decision so failures can be diagnosed.
  6. Measure your own success rate on 50–100 real tasks before reducing oversight. Benchmarks tell you the trend; only your tasks tell you your reliability.

Related: best AI agent control planes, best tools for running multiple AI coding agents, AWS Bedrock AgentCore alternatives.

Last verified: October 6, 2026. Time horizons converted from METR’s published minutes and rounded.

Sources