AI agents · OpenClaw · self-hosting · automation

Quick Answer

How to Monitor and Evaluate AI Agents in Production (2026)

Published:

The short answer

Monitoring an agent means watching outcomes, not requests. A single model call can succeed while the agent fails the task three tool calls later. As of October 2026 the working recipe is:

  1. Trace every run as a tree of spans: model calls, tool calls, retrievals, with tokens, cost and latency.
  2. Score a sample of live traces automatically (online evals) for task success, faithfulness and policy.
  3. Alert on aggregate metrics — success rate, cost per task, tool error rate — per agent version.
  4. Review low-scoring traces by hand and label them.
  5. Regress: every labelled failure becomes a test case that runs before the next deploy.

Step 1: Instrument with traces

Use OpenTelemetry-compatible tracing so you are not locked to one vendor. Most frameworks emit traces natively (LangGraph, OpenAI Agents SDK, Mastra, Spring AI via Micrometer, Koog via OpenTelemetry). Capture for each span: input, output, model and version, tokens, cost, latency, tool name and arguments, and a session_id and user_id so you can follow a conversation.

Redact secrets and personal data before export, not in the dashboard.

Step 2: Define what “success” means

Write it down per agent before you measure anything:

AgentSuccess signal
Support agentTicket resolved without human handoff; no reopen in 7 days
Coding agentPR merged; CI green; no revert in 14 days
Research agentAnswer cites sources that support each claim
Booking agentBooking created with correct fields in the system of record

Hard signals from your own systems (merged, resolved, booked) beat any judge model.

Step 3: Run online evaluations

Score a sample — 5–20% of traffic, or 100% of high-risk flows — with:

  • Code checks: valid JSON, required tool called, no forbidden tool, cost under budget.
  • LLM-as-a-judge against a rubric: faithfulness to retrieved context, task completion, tone. A decision model can return a calibrated score far cheaper than a full LLM critique.
  • User feedback: thumbs, edits, retries.

Step 4: Alert on the right metrics

MetricWhy it mattersTypical alert
Task success rateThe number your users feelDrop of 5 points vs 7-day baseline
Cost per completed taskAgents loop; tokens compound+30% vs baseline
Steps / tool calls per taskEarly sign of loops or confusionp95 doubles
Tool error rateBroken APIs, bad arguments>2%
Escalation rateHidden failure mode+50%
p95 latencyTimeouts upstreamAbove SLA

Tag every trace with the prompt, model and code version so a regression points to its cause.

Step 5: Close the loop with regression tests

Export failed traces into a dataset, label the expected outcome, and run the dataset as an offline eval in CI on every prompt, model or tool change. Gate deploys on it. This is the step most teams skip and the one that stops the same failure shipping twice.

Which tool to use

LangfuseLangSmithBraintrust
Best forOpen source, self-hosting, framework-neutralLangChain / LangGraph teamsEval-first teams running experiments
Free tier (Oct 2026)Hobby: 50k units/mo, 30 days data, 2 usersDeveloper: 5k base traces/mo, 1 seatStarter: $0, 1 GB processed data, 10k scores/mo
Paid entryCore $29/mo (100k units, then $8 per 100k); Pro $199/moPlus $39/seat/mo (10k traces), then pay as you goPro $249/mo (5 GB, 50k scores)
Self-hostYes (open source)Enterprise onlyEnterprise
Online evalsYesYesYes

For a broader ranking including Arize Phoenix, Helicone and Datadog, see best LLM observability tools ranked. For the difference between guardrails, evals and observability, see AI guardrails vs evals vs observability.

Common mistakes

  • Monitoring only latency and errors. An agent can return HTTP 200 with a wrong answer.
  • Judging with the same model that produced the answer. Use a different model family or code checks.
  • No version tags. Without them you cannot tell which change broke success rate.
  • Sampling only happy paths. Always score 100% of escalations and user complaints.

Building the agent itself? Start with how to build production AI agents.

Last verified: October 11, 2026.

Sources