AI Engineering

AI Agent Observability in Production

Dipankar Sarkar · · 6 min read

An AI agent that fails silently in production is worse than one that fails loudly, because the silent failure is the one nobody investigates until a customer complains. Observability for agents means three specific things: tracing every decision the agent made and why, logging every tool call and its result, and alerting on the handful of signals that actually predict a bad outcome — not the dozens of metrics that are easy to collect but tell you nothing actionable.

Why Agent Observability Is Different From Application Observability

Traditional application observability answers “is the service up, and how fast is it responding.” Those questions still matter for an agent, but they miss the failure modes that are specific to agentic systems: an agent that is up, fast, and confidently wrong. A slow database query shows up immediately in latency graphs. An agent that picked the wrong tool, misread a document, or looped on a task it could not complete shows up nowhere in a standard APM dashboard, because from the infrastructure’s point of view the request completed successfully — it just did the wrong thing.

This is the core problem agent observability has to solve: correctness is not visible from the outside. You need to instrument the reasoning path itself, not just the request-response cycle around it.

The Three Layers That Matter

1. Decision Tracing

Every step an agent takes — which tool it chose, what input it passed, what it decided to do with the output — needs to be captured as a structured trace, not a free-text log line. A trace lets you reconstruct exactly why the agent did what it did, which is the only way to debug a bad outcome after the fact instead of trying to reproduce it. OpenTelemetry’s tracing model works for this if you treat each agent step as a span with structured attributes (tool name, input hash, output summary, latency, and a success/failure flag), even though OpenTelemetry was not originally designed with agents in mind.

2. Tool-Call Logging

Every external call an agent makes — a database query, an API request, a file write — is a place where the agent’s action has a real-world side effect. These calls need to be logged with enough detail to answer “what did the agent actually do to the world,” separately from “what did the agent think it was doing.” The two diverge more often than teams expect: a tool call that times out, retries, and partially succeeds can leave an agent believing an action completed when it only half-completed, and only the tool-call log — not the agent’s own narration — tells you the truth.

3. Outcome Signals

Tracing and logging tell you what happened; outcome signals tell you whether it mattered. The signals worth alerting on are the ones that correlate with real failures: override rate (how often a human overrides or reverses what the agent did), escalation rate (how often the agent hands off to a human instead of completing the task), and repeated-tool-call rate (how often the agent calls the same tool with similar inputs in a short window, which usually means it is stuck). Generic metrics like token count or response time rarely predict whether the agent got the task right.

A Comparison of Observability Approaches

ApproachWhat it capturesWhat it missesBest for
Application logs onlyErrors, exceptions, request/responseThe agent’s reasoning pathDebugging crashes, not bad decisions
Full LLM call loggingEvery prompt and completionWhether the agent used a tool correctlyPrompt debugging, not system behavior
Structured decision tracingTool choice, input, output, per-step outcomeNothing structural if instrumented wellReconstructing why an agent did something
Outcome-signal monitoringOverride rate, escalation rate, loop detectionThe specific step that went wrongAlerting on trends before they become incidents

The practical answer is not to pick one row — it is to run tracing and outcome-signal monitoring together. Tracing without outcome signals means you can reconstruct any single failure but have no early warning system. Outcome signals without tracing tell you something is wrong but not what, which turns every incident into an open-ended investigation.

A Worked Example: Debugging a Stuck Agent

Consider an agent handling customer support tickets that occasionally spends far longer than expected on a ticket before either resolving it or escalating. With decision tracing in place, the debugging path looks like this:

  1. Check the outcome signal. The repeated-tool-call rate spiked for this ticket category starting three days ago — that is the alert that should have fired.
  2. Pull the trace for a representative slow ticket. The trace shows the agent calling a “search knowledge base” tool six times with slightly reworded queries before escalating.
  3. Check the tool-call log for that tool. The knowledge base search started returning fewer results per query three days ago — a search index update degraded recall.
  4. Confirm the root cause. The agent was not malfunctioning; the tool it depended on degraded, and the agent’s retry behavior (which looked reasonable in isolation) amplified the problem into a visible slowdown.

Without tracing, this would have looked like “the agent got worse” with no path to a root cause. With it, the fix was a one-line rollback of a search index change, found in under twenty minutes.

Limitations

Full decision tracing has a real cost: storage for traces grows quickly on a high-volume agent, and the engineering time to instrument every tool call properly is not trivial, especially retrofitting it onto an agent that was not built with observability in mind. Sampling helps for high-volume, low-risk agents, but sampling is the wrong choice for anything with compliance or safety requirements, where you need the trace for the one interaction that mattered, not a representative sample of interactions that did not. There is also no substitute for deciding, up front, which outcome signals actually matter for your specific agent — a generic dashboard of “agent health metrics” without that decision tends to accumulate noise nobody looks at.

FAQ

Do I need a dedicated observability platform for AI agents?

Not necessarily at first. OpenTelemetry with a standard backend (Grafana, Honeycomb, Datadog) can carry agent traces if you define the span structure yourself. Purpose-built tools like Langfuse or LangSmith add agent-specific views on top of that, which are worth adopting once the volume of traces makes a generic dashboard hard to navigate.

What is the minimum viable observability setup for a new agent?

Structured decision tracing on every step, tool-call logging with input/output, and one outcome signal — override or escalation rate — with an alert threshold. That covers the three layers above at the smallest scope that is still useful.

How is this different from evaluating an agent before deployment?

Pre-deployment evaluation tests the agent against known scenarios before it ships. Observability is what tells you the agent is behaving correctly against scenarios nobody wrote a test for, which is most of what actually happens in production. The two are complementary, not substitutes for each other.

Bottom Line

Observability for AI agents is not application monitoring with an LLM bolted on — it requires tracing the reasoning path, logging the real-world side effects separately from the agent’s own account of them, and alerting on the small number of outcome signals that actually predict failure. Building this layer is part of what turns an agent from a working demo into a system a team can trust to run unattended, which is the same problem forward deployment engineering and the Substrate Pattern are built to solve from the safety side. If you are building or hardening an agent runtime, the AI agent infrastructure practice covers observability as part of the production stack.

Dipankar Sarkar

Dipankar Sarkar

AI Enablement, Contract AI Engineering & Delivery

Related Articles