Agentic AI monitoring must follow an entire goal from request to outcome—not just watch the model API. A production design should create one trace for each agent run, connect every model call, retrieval, tool execution, handoff, guardrail, approval, and final result, and then measure reliability, latency, cost, quality, and policy compliance from that shared context.
The minimum viable operating model is straightforward: instrument the full path, define a small telemetry contract, build service-level objectives around user and business outcomes, evaluate sampled traces for quality and safety, and protect prompt and tool content as sensitive data. Infrastructure health still matters, but CPU and API latency cannot tell you whether an agent completed the task, chose the right tool, respected an approval boundary, or spent ten times the normal amount to produce a weak answer.
Operating principle: metrics tell you that the fleet changed, traces show where the run changed, evaluations indicate whether the result remained useful, and policy events prove whether controls were applied.
Monitoring, observability, and evaluation are different
Teams often use these terms interchangeably, which leads to dashboards that collect a great deal of data without answering operational questions.
| Capability | Question it answers | Typical signals | Primary use |
|---|---|---|---|
| Monitoring | Is a known condition outside its expected range? | Rates, errors, latency, saturation, token use, cost, queue depth | Alerting and service-level reporting |
| Observability | Why did this run behave differently? | Correlated traces, structured logs, events, model and tool metadata | Investigation and root-cause analysis |
| Evaluation | Was the answer or trajectory good enough? | Task completion, groundedness, relevance, policy checks, human feedback | Quality, safety, and regression detection |
| Audit evidence | What was authorized, executed, approved, and retained? | Identity, policy version, decision, approval, action result, evidence pointer | Governance, security, and forensic reconstruction |
A strong agentic AI operations program connects all four. An alert should lead to the affected trace; the trace should include evaluation and policy outcomes; and the investigation should be able to reach the governed evidence without copying sensitive payloads into every monitoring tool.
A practical agentic AI observability architecture
The architecture begins at the task boundary. Create a run or trace identifier before the agent makes its first model call, then propagate that context through the orchestrator, model gateway, retrieval layer, memory service, tool broker, nested agents, approval workflow, and destination systems. Export the resulting spans, metrics, logs, and events through a controlled telemetry pipeline.

The logical flow should include five control points:
- Instrumentation: the agent framework and application emit trace context for runs, model calls, tools, retrieval, handoffs, guardrails, and approvals.
- Collection: an OpenTelemetry-compatible collector receives, enriches, redacts, samples, and routes telemetry.
- Storage: metrics, traces, logs, evaluation results, and protected evidence use retention and access controls appropriate to their sensitivity.
- Analysis: dashboards, service-level objectives, anomaly rules, and evaluations identify operational and quality regressions.
- Response: alerts route to an owner who can pause a workflow, constrain tools, roll back a version, or invoke the organization’s incident process.
The observability backend can change without changing this logical design. Keep the telemetry contract separate from vendor-specific dashboards and query syntax.
Trace the complete agent run
The useful unit of analysis is the complete run or task trajectory. A trace should make parent-child relationships visible so an operator can distinguish model delay from retrieval delay, a failed tool from a bad plan, or an approval wait from an infrastructure problem.
agent.run
├── agent.plan
├── retrieval.search
├── model.chat
├── tool.execute
│ └── dependency.request
├── guardrail.evaluate
├── human.approval
├── agent.handoff
│ └── nested-agent.run
└── outcome.record
The OpenTelemetry GenAI agent conventions define operations such as invoke_agent, plan, execute_tool, retrieval, and memory activity. They also define attributes for provider, model, agent, conversation, token use, and finish reasons. The agent conventions are still marked as development, so pin the convention version used by your instrumentation and test schema changes before promoting them.
Use low-cardinality dimensions for metrics and dashboards: environment, service, agent name, agent version, workflow, model, tool class, outcome, and bounded error type. Keep high-cardinality values such as trace IDs, conversation IDs, task IDs, and document IDs in traces or events rather than unrestricted metric labels.
Define a minimum telemetry contract
A telemetry contract keeps frameworks and vendors from emitting incompatible records. The following fields are a practical starting point.
| Layer | Capture by default | Why it matters |
|---|---|---|
| Run and session | Trace ID, workflow, environment, start/end time, outcome, requesting service, approved tenant-safe correlation ID | Connects the complete request and supports task-level SLOs |
| Agent configuration | Agent name and version, prompt or policy version, tool-catalog version, deployment version | Shows which change introduced a regression |
| Model call | Provider, requested and response model, duration, input/output tokens, finish reason, bounded error type | Separates model performance and cost from the rest of the workflow |
| Tool execution | Tool name and version, action class, duration, result class, retries, policy decision, side-effect class | Finds dependency failures, unsafe calls, and retry loops |
| Retrieval and memory | Data-source class, result count, duration, cache result, access decision, retrieval-quality signal | Identifies missing context, stale memory, and slow data paths |
| Guardrail and approval | Control name and version, allow/deny/constrain result, approval state, enforcement confirmation | Proves that a policy decision became an actual control |
| Outcome | Task completion, user correction, escalation, abandonment, deterministic checks, sampled evaluation scores | Connects technical behavior to user and business value |
Record full prompts, completions, retrieved documents, and tool arguments only under an explicit content-capture policy. Metadata is enough for many operational questions and carries much less privacy and security risk.
Measure the metrics that explain agent behavior
Infrastructure metrics remain necessary, but agent operations need outcome and trajectory metrics as well. Start with a small set that changes engineering or operational decisions.
| Metric family | Examples | Operational question |
|---|---|---|
| Completion and quality | Task-success rate, completion-without-escalation rate, user correction rate, deterministic check pass rate | Did the agent produce an acceptable result? |
| Latency | End-to-end p50/p95/p99, model latency, tool latency, retrieval latency, approval wait | Where is the user waiting? |
| Reliability | Run error rate, timeout rate, dependency failures, abandoned runs, telemetry loss | Can the service complete work consistently? |
| Trajectory | Model turns per run, tool calls per success, retries, loop depth, handoffs, nested-agent fan-out | Is the workflow becoming inefficient or unstable? |
| Efficiency | Input/output tokens, cost per run, cost per successful task, cache hit rate, compute consumption | Is spend increasing faster than useful work? |
| Safety and governance | Guardrail blocks, approval requests, unauthorized-tool attempts, policy-enforcement failures, human overrides | Are controls being invoked and enforced? |
Cost per request can be misleading because a cheaper failed run has no value. Pair cost with task outcome and workflow version.
task_success_rate = successful_runs / completed_runs
cost_per_success = total_model_and_tool_cost / successful_runs
tool_retry_rate = repeated_tool_calls / tool_calls
guardrail_enforcement_rate = confirmed_blocks / deny_decisions
telemetry_completeness = runs_with_required_spans / completed_runs
These are design formulas, not vendor-specific queries. Define the exact numerator, denominator, exclusions, and time window in the service-level objective so different teams calculate the same result.
Build service-level objectives and alerts around failure modes
Alert on user impact, control failure, and sustained deviation—not every unusual trace. Thresholds depend on the workflow’s risk and traffic, but the response logic can be standardized.
| Trigger | Likely question | First response |
|---|---|---|
| Task-success rate drops after a deployment | Did the prompt, model, tool catalog, or policy version change? | Compare versions and sampled failed traces; roll back when the regression is clear |
| End-to-end p95 rises while model latency is stable | Is a tool, retrieval service, queue, or approval path slow? | Use child-span duration to identify the dependency |
| Tool retries or repeated calls increase | Is the agent stuck, receiving ambiguous output, or ignoring a stop condition? | Constrain retry and action budgets; inspect the repeated trajectory |
| Cost per successful task rises | Did token use, model choice, failed attempts, or cache behavior change? | Break down cost by agent, version, model, and outcome |
| Deny decisions lack confirmed enforcement | Did the guardrail report a block without stopping the action? | Treat as a control failure and pause high-risk execution |
| Required spans disappear while runs continue | Is collection broken or is a component bypassing instrumentation? | Restore visibility; fail closed for workflows whose controls depend on telemetry |
| Online evaluation pass rate degrades | Is quality changing even though the service remains healthy? | Review representative traces and reproduce them in an offline evaluation set |
Use burn-rate or sustained-window alerts for service-level objectives when traffic is high enough. For low-volume, high-impact agents, deterministic policy failures and destructive-action anomalies may need immediate paging even when aggregate rates look normal.
Quality and safety evaluations need their own feedback loop
A trace can be complete and fast while the answer is wrong. Pair operational telemetry with three evaluation layers:
- Deterministic checks: schema validity, required citations, allowed tools, expected records, policy decisions, and task-specific business rules.
- Sampled online evaluations: relevance, groundedness, refusal quality, trajectory quality, and goal completion on production traces.
- Human review: high-risk cases, low-confidence results, disagreements between evaluators, and samples used to calibrate automated scoring.
Run important scenarios offline before deployment, then move a controlled subset of checks into production. The LangSmith evaluation workflow distinguishes offline testing from online monitoring, while Datadog’s trace-level evaluation guidance shows why a multi-step agent often must be judged across the whole trace rather than one model span. Treat model-based evaluators as signals to calibrate and review, not unquestionable ground truth.
Protect prompts, tool data, and user context
Agent telemetry can contain credentials, personal data, customer records, source code, security findings, retrieved documents, system instructions, and tool results. Content capture should therefore be an explicit architecture decision.
- Capture identifiers, versions, timings, result classes, token counts, and policy decisions by default.
- Redact secrets and regulated fields before export rather than relying only on access controls after storage.
- Use hashes or secure evidence pointers when operators need correlation but not the raw payload.
- Separate tenant data and enforce least-privilege access to traces, dashboards, and exports.
- Define retention by signal and risk. Routine metrics, full content, security events, and incident evidence should not share one retention policy.
- Make dropped spans, sampling decisions, redaction failures, and collector backpressure observable.
This is consistent with current platform behavior. OpenTelemetry treats message and system-instruction content as opt-in, and the OpenAI Agents SDK tracing controls allow sensitive model and tool inputs and outputs to be excluded while retaining spans. Apply the same principle regardless of framework: collect the least content needed for the operational and governance objective.
Use a vendor-neutral instrumentation pattern
Use framework instrumentation when it emits the context you need, then add application spans for business outcomes, approvals, or custom tools that the framework cannot see. The example below intentionally uses custom app.* attributes for organization-specific fields and avoids raw prompt content.
from opentelemetry import trace
tracer = trace.get_tracer("agent-runtime")
def run_agent(task, agent):
with tracer.start_as_current_span("agent.run") as run_span:
run_span.set_attribute("app.workflow.name", task.workflow)
run_span.set_attribute("app.agent.name", agent.name)
run_span.set_attribute("app.agent.version", agent.version)
run_span.set_attribute("app.environment", task.environment)
try:
with tracer.start_as_current_span("tool.execute") as tool_span:
tool_span.set_attribute("app.tool.name", task.tool_name)
tool_span.set_attribute("app.tool.side_effect", task.side_effect_class)
result = execute_approved_tool(task)
run_span.set_attribute("app.outcome", "success")
return result
except Exception as exc:
run_span.record_exception(exc)
run_span.set_attribute("app.outcome", "error")
raise
Adapt the example to your runtime, propagate context through asynchronous workers and remote tools, and use current gen_ai.* conventions for model and agent operations when your instrumentation supports them. A useful validation test starts one synthetic run and confirms that every expected child operation appears under the same root trace with the correct agent and deployment version.
Design dashboards for decisions, not data volume
One giant dashboard usually satisfies no audience. Build a small set of views with explicit owners.
- Executive and product view: successful tasks, adoption, escalation, cost per success, and business outcome.
- Operations view: traffic, end-to-end latency, error and timeout rates, queue depth, tool reliability, and current SLO burn.
- Engineering view: trace waterfall, model and tool breakdown, retries, version comparison, token use, cache behavior, and failed examples.
- Quality view: deterministic checks, sampled evaluation scores, human feedback, failure taxonomy, and regression by version.
- Governance and security view: guardrail decisions, approval outcomes, policy-enforcement confirmation, sensitive-content capture, and audit completeness.
Every aggregate should link to representative traces. An operator should be able to move from a failed SLO to a version, a dependency, a trace, and an accountable owner without switching among uncorrelated identifiers.
Create an incident workflow for agent failures
- Classify impact: user delay, incorrect output, policy failure, unauthorized action, cost spike, or telemetry loss.
- Identify the affected version: agent, prompt, model, tool catalog, policy, and deployment.
- Open representative traces: compare successful and failed runs with the same workflow and environment.
- Contain the problem: pause the workflow, remove a tool, constrain budgets, route to a human, or roll back the version.
- Verify recovery: use synthetic runs and production SLOs to confirm that both service and telemetry have recovered.
- Turn evidence into prevention: add failed traces to the offline evaluation set and update the runbook, alert, or policy test.
For high-authority agents, monitoring must connect to independent controls. The AI gateway operating model explains how identity, policy, observability, and cost controls fit at the gateway layer.
A 30-day rollout plan
| Period | Deliverable | Exit criterion |
|---|---|---|
| Days 1–7 | Select one agent workflow; define run outcomes, owners, versions, and required spans | A synthetic run produces one complete trace with no sensitive content captured unintentionally |
| Days 8–14 | Add completion, latency, reliability, tool, token, cost, and telemetry-health metrics | Dashboards distinguish model, retrieval, tool, approval, and end-to-end time |
| Days 15–21 | Define initial SLOs, alert routes, deterministic checks, and a small offline evaluation set | Each alert has an owner, runbook, representative trace, and safe containment action |
| Days 22–30 | Enable sampled online evaluations, privacy review, retention policy, and failure exercises | The team can detect, investigate, contain, and reproduce a controlled failure |
Do not begin by instrumenting every agent. Start with a workflow that has real users, meaningful business value, observable outcomes, and an owner who can act on the data.
Where advanced divergence detection begins
This guide owns the production monitoring foundation: complete traces, meaningful metrics, quality evaluations, SLOs, dashboards, and privacy controls. High-authority agents also need trajectory-level security that detects privilege expansion, unrelated destinations, credential discovery, persistence, lateral movement, and failed intervention. Continue with Agent Observability Is Not Logging: How to Detect Autonomous System Divergence in Real Time for that deeper control model.
Bottom line
Agentic AI monitoring succeeds when an operator can answer five questions quickly: What goal was the agent pursuing? Which version and dependencies executed? Where did the run slow down or fail? Was the outcome useful and policy-compliant? What action should the team take now?
Build one correlated trace per run, standardize the minimum telemetry contract, measure cost and quality per successful outcome, treat content capture as sensitive, and connect every alert to a trace, owner, and response. That creates an operating system for agents instead of another pile of logs.
Official and primary references
- OpenTelemetry: GenAI observability with OpenTelemetry
- OpenTelemetry: AI agent observability standards and practices
- OpenTelemetry: GenAI agent and framework span conventions
- OpenAI Agents SDK: tracing
- Microsoft Agent Framework: workflow observability
- Datadog: Agent Observability metrics
- LangSmith: observability concepts
- NIST AI 600-1: Generative AI Profile