The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →A wrong answer from an AI application is a symptom, not a diagnosis. “Where did it break?” invites you to blame the model, because the model is the most visible part. “Which layer diverged from expected behavior?” sends you to the earliest step in the request path where reality departed from the plan. That step might be the prompt, retrieval, a tool call, the model, or the infrastructure underneath. This guide shows how to find it.
The layers below are a working model, not a universal fixed stack. Real systems blend or skip them, and the sources cited here describe provider-specific tooling.
The five failure classes to separate
AWS’s guidance on improving generative AI applications makes a distinction many teams skip. The model and knowledge base may be fully capable and still produce poor output because the application gave them the wrong instructions. AWS describes this as a software-layer problem. Building on that, keep these classes apart:
1. Prompt and orchestration
The prompt template is poor, the request is routed to the wrong path, or the wrong tool or agent action is chosen. The model may have done exactly what it was told.
#1 Best Overall
2. Knowledge and retrieval
The needed information is missing, stale, incorrect, inaccessible, or simply not retrieved. In a retrieval-augmented generation (RAG) flow, check what context actually reached the model. Don’t assume it saw the source material you expected. (AWS guidance)
3. Core model
Given suitable instructions and good context, the foundation model may still lack the specialized knowledge, reasoning ability, or stylistic range the task needs. This is the correct diagnosis only after the layers above have been cleared.
Rank #2
4. Tool and external-service execution
Agents act through tools and APIs. Google’s agent observability documentation lists tool usage, call counts, success or failure, latency, and exchanged data as things you can observe. A correct tool choice followed by a failed API call is a different problem from a successful call that returned unsuitable data.
5. Application and infrastructure
Errors and latency can originate in application code, data stores, or supporting services. Google’s AI and ML reliability guidance (last reviewed 2025-08-07) recommends observability across infrastructure, application code, data, and model behavior together. AWS’s CloudWatch documentation likewise treats generative AI applications alongside their underlying infrastructure.
Recommended Free Tools
These classes overlap. A stale document is both a knowledge problem and, possibly, an ingestion-pipeline problem. The aim is not tidy labels but the earliest divergence point, followed by a test of whether changing that step fixes the failure.
An investigation sequence that works for most failures
- Capture the failing case. Record the user input, time, environment, application, model and configuration versions, and the outcome you expected. Keep identifiers so the interaction can be found again.
- Follow one trace end to end. Look at the prompt and routing decision, retrieved context, model request and response, tool calls, post-processing, and final reply. CloudWatch documents end-to-end prompt traces across knowledge bases, tools, and models. Google describes traces as execution paths that can expose model calls and tool use.
- Check the inputs at each boundary. Verify what was actually supplied: instructions, retrieved passages, permissions, tool arguments, service responses. For RAG, ask whether the right material existed and whether it was retrieved. Google names context relevance and response groundedness as monitoring concerns.
- Correlate logs and metrics. Use a trace or interaction ID to pull the matching logs and service signals. AWS recommends structured logs, trace IDs, and custom metrics by layer to help separate model-related errors from infrastructure problems.
- Compare against a baseline. Look at correctness and groundedness next to latency, errors, throttling, token use, retrieval relevance, and tool success and latency. CloudWatch’s documented metrics include invocation totals, token usage, latency percentiles, errors, throttling, and cost attribution.
- Change one plausible cause and re-evaluate. Match the fix to the layer, as the table shows.
Matching the symptom to the layer and the fix
| What the trace shows | Likely layer | First fix to try |
|---|---|---|
| Wrong subagent, route, or action chosen | Prompt and orchestration | Adjust agent or prompt configuration and instructions |
| Right answer absent from retrieved passages | Knowledge and retrieval | Fix ingestion, access, ranking, or the source corpus |
| Right passages retrieved, answer ignores or contradicts them | Prompt or core model | Tighten grounding instructions; if that fails, test another model |
| Correct tool chosen, call errored or timed out | Tool or external service | Inspect request, response, error, and latency of that call |
| Tool succeeded but data was unsuitable | Tool design or data source | Review what the tool returns and how it is described to the agent |
| Inputs and context look sound, task exceeds capability | Core model | Try a more suitable model, decompose the task, or add human review |
| Slow or failing responses with throttling or errors | Application or infrastructure | Follow errors and latency through application code and supporting services |
The fixes in the last column are operational recommendations derived from this method. The cited documentation does not measure their effectiveness. Keep representative failures as evaluation cases so you can confirm a fix and catch regressions later.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.RAG and agents: a worked example of following execution order
Salesforce’s troubleshooting guide for knowledge retrieval in agents starts at the agent layer. First verify that the correct subagent and action were selected and executed. Then inspect the agent instructions and action instructions. Only after that does it move to the data library: check its status and permissions, then inspect the indexed chunks and the retrieval results. The ordering is the lesson. You work through the path the request took, and you don’t jump to the model.
For any agent, split the tool question in two: why was this tool chosen, and what did it return? Google recommends monitoring tool calls, outcomes, latency, and exchanged data, which gives you the evidence for each half.
Best Value
Three signals, three jobs
- Traces show the execution path and order of steps.
- Logs preserve event and error detail for a given step.
- Metrics track rates, latency, and usage over time, so you can tell a one-off from a trend.
A shared trace or interaction ID is what lets you move between the three without guessing.
Choosing observability tooling
No page reviewed here gives comparable prices or a complete feature matrix across vendors, so a “best tool” ranking would be unsupported. Evaluate any option on these axes instead:
- Coverage of model, retrieval, agent and tool, application, and infrastructure components.
- Whether traces expose intermediate inputs, outputs, and execution order.
- Metrics for latency, errors, token use, retrieval, and tool outcomes.
- Correlation of traces with structured logs and alerts.
- Framework and provider compatibility, data-handling controls, and operational cost.
AWS and Google both document first-party capabilities in this area (CloudWatch, Google Cloud Observability). Confirm what your own stack’s provider supports.
The Bottom Line
Treat every bad answer as the end of a chain. Trace it back to the first step where what happened differs from what should have happened, and change only that step. Blame the model last, after you’ve shown that the instructions, context, tools, and infrastructure were sound.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




