Blog
OpenTelemetry for AI Agents: Tracing Reasoning Loops, Tool Chains, and Token Costs
Traditional APM tells you an HTTP request returned 200 OK, but says nothing about whether your agent entered a circular retry loop or hallucinated a file path. Standardizing on OpenTelemetry GenAI semantic conventions turns opaque reasoning cycles into queryable traces.
OpenTelemetry for AI agents extends distributed tracing beyond stateless HTTP requests to model non-deterministic, multi-step execution loops. By adopting OpenTelemetry GenAI semantic conventions alongside agent-specific spans, teams capture root-level goals, inner reasoning cycles, tool invocations, and subagent delegations in a single directed acyclic graph. Each span records standard attributes such as gen_ai.provider.name, gen_ai.request.model, input and output token counts, tool arguments, and failure reasons. This structured telemetry replaces unsearchable terminal scrollback with queryable traces, enabling developers to identify circular retries, runaway token spend, and broken decision branches across autonomous agent squads in production.
Why traditional APM fails for autonomous agents
Traditional application performance monitoring assumes a linear, deterministic request-response lifecycle. A client sends an HTTP request, a server executes predictable business logic, queries a database, and returns a response with a status code. If the server responds with 200 OK, monitoring systems record a successful transaction. If it responds with 500 Internal Server Error, alerting rules fire.
Autonomous coding agents violate every assumption of that model. An agent task is an iterative, open-ended goal loop that can execute dozens of LLM completions, run multiple shell commands, and edit several files over twenty minutes. Every individual HTTP request to an inference provider returns 200 OK, yet the agent may be caught in a circular retry loop, modifying the wrong directory, or hallucinating a non-existent build tool.
When an agent fails, the failure is rarely a crash or a network timeout. The failure is logical: the agent misinterprets a compiler error, invokes an inappropriate tool, or loses track of its original objective as the context window fills up. Searching through megabytes of raw terminal scrollback across multiple parallel agents is unworkable. Observability for agents requires tracking the decision tree itself rather than surface-level HTTP codes.
The OpenTelemetry GenAI semantic conventions
OpenTelemetry provides a vendor-neutral standard for traces, metrics, and logs across distributed architectures. Rather than relying on proprietary logging formats from individual model vendors, the OpenTelemetry community established official GenAI semantic conventions to standardize LLM telemetry.
At the foundation of these conventions are uniform span attributes. Every call to an LLM provider emits a client span capturing the engine name, model identity, token usage, and inference parameters. Standardizing these keys ensures that telemetry ingested into any OTel-compatible backend, such as Jaeger, Honeycomb, or Datadog, can be aggregated and analyzed across different model vendors without custom translators.
- gen_ai.provider.name: Identifies the model provider or API, such as anthropic, openai, or google. It replaced the deprecated gen_ai.system.
- gen_ai.request.model: The specific model version requested, such as claude-3-7-sonnet or gpt-4o.
- gen_ai.usage.input_tokens: The number of input tokens consumed by the prompt and history. It replaced the deprecated gen_ai.usage.prompt_tokens.
- gen_ai.usage.output_tokens: The number of tokens generated in the model's response. It replaced the deprecated gen_ai.usage.completion_tokens.
- gen_ai.response.finish_reasons: Why the generation stopped, such as stop, tool_calls, or length.
Agent-specific tracing: mapping the execution hierarchy
Standard GenAI spans capture individual LLM calls, but an agent is an orchestrator of loops rather than a single invocation. Full agent observability requires a hierarchical trace structure that mirrors the agent's actual runtime execution.
The root span of a trace represents the user's high-level task or objective. Beneath the root span sit child spans representing individual reasoning iterations. Each reasoning cycle contains an LLM generation span followed by one or more tool execution spans. When an agent invokes a shell command, queries a database, or calls a Model Context Protocol tool, the tool invocation runs as a child span linked to the reasoning step that produced it.
In multi-agent architectures, this hierarchy extends across processes. When a primary coordinator agent delegates a subtask to a specialized worker or research subagent, it injects W3C trace context into the delegation payload. The subagent's execution loop becomes a nested child span within the parent's distributed trace. Developers can expand the trace to follow the delegation chain from high-level architectural planning down to a specific bash command.
Debugging decision trees: capturing why, not just what
A traditional process monitor records that an agent ran git checkout -b feature-branch and exited with code 0. That confirms execution, but it does not reveal why the agent chose that command, what alternatives it rejected, or how it reacted to the output.
Effective agent tracing captures the decision context alongside the execution payload. Each reasoning span should record the agent's chain-of-thought, the system prompt state, the tool definitions available at that step, and the explicit tool call arguments selected by the model. When a tool fails, the subsequent span must capture the agent's reflection: did it understand the error output, or did it repeat the identical command with identical arguments?
Capturing this context introduces a critical data sanitization requirement. Prompts and tool parameters frequently encounter API tokens, private repository URLs, and database credentials. An observability pipeline must redact secrets at the telemetry collector boundary before spans leave the local machine. If raw credentials land in persistent trace storage, distributed tracing becomes an exfiltration vector.
Squad and production observability: metrics over scrollback
Running a single agent in a local terminal allows a developer to watch the scrollback in real time. Running a squad of five autonomous agents across multiple repositories makes manual inspection impossible. Reading megabytes of raw console logs across concurrent agent runs leads to prompt fatigue and missed failures.
Production agent observability shifts supervision from raw text inspection to aggregated telemetry metrics. By extracting metrics from OpenTelemetry spans, engineering leads can monitor squad health through four primary indicators: completion velocity, tool failure rates, token consumption per task, and reasoning loop depth.
When an agent exceeds a threshold of ten consecutive reasoning loops without completing a task milestone, alerting rules flag the trace as anomalous. Developers can jump directly to the failing span in the trace viewer, inspect the tool inputs and model reasoning at that exact turn, and intervene before the agent burns through its context budget.
How Forkbench handles runtime observability
Forkbench approaches agent observability from the desktop runtime layer. When running autonomous coding agents like Claude Code, Codex CLI, or custom workers, developer visibility cannot depend on developers manually instrumenting proprietary agent source code.
Instead of forcing agents to rewrite their internals, Forkbench provides runtime supervision through isolated process panes, structured teamwork boards, and resource monitoring. Each pane carries a live pulse driven by CPU and output rate, flags when its agent needs a human, and shows the board task it claimed. If an agent hangs or enters an unconstrained loop, the platform flags the stalled thread without requiring manual log audits.
Forkbench also integrates external context boundaries. A Vault secret bound to a destination is substituted into the outgoing request by a local proxy, so the agent and its subprocess only ever hold a stand-in, and every resolution is recorded for audit. A secret with no pinned destination falls back to environment injection, where the receiving process can still read it.
Related: Agent-to-Agent Protocol: distributed tracing across agents, Debug Adapter Protocol: inspecting runtime execution state, Model Context Protocol: tracing tool invocations and context, Teamwork Board: orchestrating multi-agent squads, Redacting secrets from agent traces and telemetry logs
Frequently asked
What is the difference between standard APM and AI agent tracing?
Traditional APM monitors deterministic request-response lifecycles, measuring HTTP latency, database queries, and 500-level error rates. AI agent tracing monitors non-deterministic, multi-step loops where every underlying HTTP request may return 200 OK while the agent's logical execution fails. Agent tracing models the hierarchy of goals, reasoning cycles, tool invocations, and subagent delegations defined under the OpenTelemetry GenAI Semantic Conventions.
What are the core OpenTelemetry GenAI semantic conventions?
The OpenTelemetry GenAI semantic conventions define standardized span names and attributes for generative AI operations. Core attributes include gen_ai.provider.name (the model vendor), gen_ai.request.model (the model identifier), gen_ai.usage.input_tokens (input token count), gen_ai.usage.output_tokens (output token count), and gen_ai.response.finish_reasons (completion termination cause). Using these attributes ensures interoperability across different observability backends.
How does distributed tracing work across multi-agent systems?
Multi-agent tracing can use W3C Trace Context propagation across systems like the Agent-to-Agent protocol (A2A), though A2A's own tracing conventions are still being formalized. When a coordinator agent delegates a task to a subagent or worker process, it passes trace identifiers (such as traceparent and tracestate) inside the delegation message or environment variables. The subagent initializes its spans using the parent context, linking all subagent thoughts, tool calls, and completions into a single unified trace tree.
How can teams prevent sensitive credentials from leaking into agent traces?
Agent prompts, completions, and tool inputs frequently handle sensitive data such as API keys, environment variables, and proprietary code. Teams must implement redaction filters within the OpenTelemetry SDK or OpenTelemetry Collector before exporting telemetry to external dashboards. Runtime isolation platforms like Forkbench keep credentials out of agent environment variables altogether (see Forkbench security), preventing secrets from entering trace payloads at the source.