To debug an AI agent, trace the whole workflow—not just its final answer. A useful trace links the agent run to its model generations, tool calls, handoffs, retrieval work, and application-specific steps, so you can see what ran, in what order, how long it took, and where it failed. Logs add searchable events and application context; neither logs nor traces, by themselves, prove that an answer is correct or safe.
What observability shows in an agent workflow
A user may experience one task while the system performs several operations: an agent invokes a model, calls a tool, hands work to another agent, retrieves information, or retries after an error. Looking only at the final response hides that execution path. Tracing makes related operations inspectable as one workflow. OpenAI’s Agents SDK documentation describes default trace instrumentation for model generations, function and tool calls, handoffs, guardrails, and custom events; AWS OpenSearch documentation describes hierarchical traces across orchestration, model calls, tools, and retrieval.
Trace, span, session, and turn
- Trace: a record grouping an end-to-end workflow or operation.
- Span: a record of one operation, with timing, status, and any captured attributes or content. Parent-child nesting shows how an operation fits inside its caller.
- Session and turn: some APIs add these higher-level groupings. In OpenAI’s Agents API terminology, a session can contain multiple turns, and a turn’s trace groups its steps, such as model responses, tool calls, and delegated work.
These names and exact boundaries vary by framework. Define what your own root operation represents and verify how your instrumentation nests work rather than assuming every system uses the same hierarchy. OpenAI Agents SDK tracing and the Agents API trace guide document their respective structures.
How logs fit alongside traces
Structured logs are useful for searchable events and application context; traces show how related operations fit together and where time or errors accumulate. Correlate the two with stable identifiers your application controls, such as a run ID or trace ID, so an event can be connected to its execution path. A final answer is an output, not a record of all the work that produced it.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
What to instrument
Instrument the execution path that your team needs to operate and investigate. Include meaningful operations, not every trivial internal function. A full path commonly includes:
- The agent invocation or workflow root.
- Each relevant model generation.
- Tool execution, including tool name and call identifier where available.
- Delegation or handoff between agents.
- Retrieval activity that affects the result.
- Application-specific work that materially affects the outcome but is not already represented.
For useful filtering and grouping, record stable identifiers and low-cardinality dimensions such as workflow name, provider, model, operation type, and status. OpenTelemetry’s GenAI conventions recommend meaningful, low-cardinality workflow names and caution against inventing a conversation ID when none exists: do not substitute a random UUID, trace ID, or hash of request content. Set a conversation ID only when the instrumented library already has one or your application supplies one. The conventions are a living document, so check the current guidance when implementing them: OpenTelemetry GenAI agent span conventions.
Check what automatic instrumentation actually captures
Built-in or automatic instrumentation can reduce setup, but its coverage depends on the library, provider, runtime, and configuration. AWS documents auto-instrumentation for several frameworks and providers, along with GenAI trace attributes and OpenSearch querying; that is a vendor description of its capabilities, not an independent coverage assessment. Export a sample trace and confirm that the steps you care about—including tool, retrieval, and handoff activity—appear with useful parentage, status, timing, and attributes. Add custom spans for important blind spots. AWS’s overview is at Amazon OpenSearch Service generative AI observability; an example of manual instrumentation is in OpenSearch manual GenAI instrumentation.
How to investigate a bad, failed, or slow run
- Find the run or session. Filter using identifiers your application records, then narrow to the relevant run or time window. The OpenAI Agents API trace UI documents filtering by model, status, or date and opening a session timeline.
- Follow the tree and timeline. Start at the workflow root and inspect child spans for model responses, tools, and delegated work. Find the first failed span, unexpected result, retry, or unusually long operation. The timeline helps reveal sequence, overlap, and duration.
- Inspect the relevant span. Where capture is enabled, compare model inputs and outputs or tool arguments and results. Check provider and model identifiers, tool name and call ID, status, error details, duration, and token usage where available.
- Interpret missing usage carefully. An unknown or blank usage field does not mean zero. The Agents API guide notes that usage can arrive after a turn and may change as it becomes available; it should not be treated as a final bill.
- Reproduce or isolate the operation. Use the trace to identify the operation and its surrounding context, then reproduce it with appropriately sanitized inputs or test the tool/model boundary independently.
- Fill only the remaining blind spots. Add custom spans for important application work that is not already visible, using stable names and attributes that make filtering useful.
The Agents API guide explains the session timeline, step details, and trace export: OpenAI Agents API tracing. Its session-traces endpoint returns OTLP JSON, but export must be enabled for the organization and requires suitable project permissions. SDK custom spans and processors are documented in the JavaScript Agents SDK tracing guide and Python Agents SDK tracing guide.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #3
Execution evidence is not a quality verdict
A trace can show what inputs and outputs were recorded, which operations ran, their status, and how long they took. Those facts can localize a failure or a likely source of delay. They do not establish that an answer is factually correct, policy-compliant, or safe; those judgments need appropriate evaluation and review beyond execution telemetry.
Choosing built-in tracing or OpenTelemetry
There is no universal winner established by the cited documentation. Choose according to your framework and operational needs, then inspect real exported traces before relying on coverage.
Rank #4
| Route | What it offers | What to verify |
|---|---|---|
| Framework or SDK built-in tracing | OpenAI Agents SDK documents default trace and span creation, sensitive-data settings, and export processors. JavaScript server runtimes enable tracing by default, while browsers and test mode default to disabled; Python tracing is described as enabled by default. | Confirm behavior for the exact runtime, package version, and configuration you deploy. Check which events are captured and how they are exported. |
| OpenTelemetry instrumentation and a backend | OpenTelemetry GenAI conventions provide shared attribute guidance. AWS documents OpenSearch AI traces, integration and auto-instrumentation capabilities, and querying with PPL. Manual instrumentation can add invocation and nested tool spans. | Verify instrumentor coverage and exported structure for each library/backend combination, along with export permissions and destination configuration. |
Compare options on framework and provider coverage; whether tools, retrieval, handoffs, and custom work appear; span detail; sensitive-data controls; export flexibility; correlation with logs and metrics; filtering and querying; and operational fit. The cited sources do not establish an independent product comparison, a complete vendor matrix, or a pricing comparison, so they do not support declaring one route universally better.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Protect sensitive data in traces
Trace content can include user prompts, model outputs, function inputs and results, or audio data. OpenAI’s JavaScript and Python Agents SDK documentation describes settings to disable sensitive-data capture; the Python guide states that capture is enabled by default. OpenTelemetry warns that input-message attributes can contain sensitive or personal information. Treat trace content as collected application data, not harmless diagnostics.
Best Value
- Decide which content is genuinely necessary to diagnose the system; omit or redact the rest before production.
- Configure capture controls deliberately for the runtime and SDK in use.
- Restrict access to trace content and align retention with your application’s data policy.
- Inspect an exported trace to confirm that redaction and omission work as intended.
See the JavaScript Agents SDK tracing guide, Python Agents SDK tracing guide, and OpenTelemetry GenAI conventions for the documented capture controls and content considerations.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




