How to Monitor and Evaluate AI Agents in Production (2026)
The short answer
Monitoring an agent means watching outcomes, not requests. A single model call can succeed while the agent fails the task three tool calls later. As of October 2026 the working recipe is:
- Trace every run as a tree of spans: model calls, tool calls, retrievals, with tokens, cost and latency.
- Score a sample of live traces automatically (online evals) for task success, faithfulness and policy.
- Alert on aggregate metrics — success rate, cost per task, tool error rate — per agent version.
- Review low-scoring traces by hand and label them.
- Regress: every labelled failure becomes a test case that runs before the next deploy.
Step 1: Instrument with traces
Use OpenTelemetry-compatible tracing so you are not locked to one vendor. Most frameworks emit traces natively (LangGraph, OpenAI Agents SDK, Mastra, Spring AI via Micrometer, Koog via OpenTelemetry). Capture for each span: input, output, model and version, tokens, cost, latency, tool name and arguments, and a session_id and user_id so you can follow a conversation.
Redact secrets and personal data before export, not in the dashboard.
Step 2: Define what “success” means
Write it down per agent before you measure anything:
| Agent | Success signal |
|---|---|
| Support agent | Ticket resolved without human handoff; no reopen in 7 days |
| Coding agent | PR merged; CI green; no revert in 14 days |
| Research agent | Answer cites sources that support each claim |
| Booking agent | Booking created with correct fields in the system of record |
Hard signals from your own systems (merged, resolved, booked) beat any judge model.
Step 3: Run online evaluations
Score a sample — 5–20% of traffic, or 100% of high-risk flows — with:
- Code checks: valid JSON, required tool called, no forbidden tool, cost under budget.
- LLM-as-a-judge against a rubric: faithfulness to retrieved context, task completion, tone. A decision model can return a calibrated score far cheaper than a full LLM critique.
- User feedback: thumbs, edits, retries.
Step 4: Alert on the right metrics
| Metric | Why it matters | Typical alert |
|---|---|---|
| Task success rate | The number your users feel | Drop of 5 points vs 7-day baseline |
| Cost per completed task | Agents loop; tokens compound | +30% vs baseline |
| Steps / tool calls per task | Early sign of loops or confusion | p95 doubles |
| Tool error rate | Broken APIs, bad arguments | >2% |
| Escalation rate | Hidden failure mode | +50% |
| p95 latency | Timeouts upstream | Above SLA |
Tag every trace with the prompt, model and code version so a regression points to its cause.
Step 5: Close the loop with regression tests
Export failed traces into a dataset, label the expected outcome, and run the dataset as an offline eval in CI on every prompt, model or tool change. Gate deploys on it. This is the step most teams skip and the one that stops the same failure shipping twice.
Which tool to use
| Langfuse | LangSmith | Braintrust | |
|---|---|---|---|
| Best for | Open source, self-hosting, framework-neutral | LangChain / LangGraph teams | Eval-first teams running experiments |
| Free tier (Oct 2026) | Hobby: 50k units/mo, 30 days data, 2 users | Developer: 5k base traces/mo, 1 seat | Starter: $0, 1 GB processed data, 10k scores/mo |
| Paid entry | Core $29/mo (100k units, then $8 per 100k); Pro $199/mo | Plus $39/seat/mo (10k traces), then pay as you go | Pro $249/mo (5 GB, 50k scores) |
| Self-host | Yes (open source) | Enterprise only | Enterprise |
| Online evals | Yes | Yes | Yes |
For a broader ranking including Arize Phoenix, Helicone and Datadog, see best LLM observability tools ranked. For the difference between guardrails, evals and observability, see AI guardrails vs evals vs observability.
Common mistakes
- Monitoring only latency and errors. An agent can return HTTP 200 with a wrong answer.
- Judging with the same model that produced the answer. Use a different model family or code checks.
- No version tags. Without them you cannot tell which change broke success rate.
- Sampling only happy paths. Always score 100% of escalations and user complaints.
Building the agent itself? Start with how to build production AI agents.
Last verified: October 11, 2026.