Retrieval-augmented generation (RAG) lets an AI application look up relevant information in an external collection and pass selected material to a large language model (LLM) before it answers. That can help the model respond using information such as an organization’s documents, but RAG does not guarantee that the answer is correct: the data, search, prompt, and model response all matter.
What is RAG in AI?
RAG combines information retrieval—the process of searching a collection for useful material—with a language model that generates a response. Instead of relying only on information encoded during training, a RAG application searches an external source at answer time, places selected results in the model’s prompt, and asks the model to use that context.
AWS Prescriptive Guidance defines it this way: “Retrieval Augmented Generation (RAG) is a technique used to augment a large language model (LLM) with external data, such as a company’s internal documents.” See AWS Prescriptive Guidance: Understanding Retrieval Augmented Generation.
The external collection might contain proprietary or otherwise task-specific information. RAG is an architecture pattern, not a guarantee of factuality or a single product: what it can answer depends on what has been prepared and indexed, what the search retrieves, and how the model uses the supplied context.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
How does retrieval-augmented generation work?
A basic system has two connected paths: a preparation path that makes source material searchable, and a query-time path that retrieves material for a particular question. Microsoft’s RAG solution design and evaluation guide describes the stages below.
1. Prepare and index the source data
Documents or other supported media enter a data pipeline. The system divides them into chunks—smaller passages intended to preserve useful meaning—and may attach metadata such as titles or summaries. When vector search is used, it creates embeddings, numerical representations that help identify material with related meaning. Processed material is then stored in a search index.
These choices affect what can be found later. Chunks that are too broad may include distracting material; chunks that are too narrow may lose necessary context. Metadata and indexing also shape which passages a search can return.
Rank #2
2. Receive the user’s question
The application sends the question to an orchestrator, the part of the system that coordinates search and the model call. It may pass the question as written or prepare it for the configured search process.
Recommended Free Tools
3. Search for supporting material
The orchestrator searches the configured collection and selects results. Search may use vector retrieval, full-text search, a hybrid of both, or multiple searches in sequence. These methods are not interchangeable defaults: the right choice depends on the material and the questions people ask.
4. Build context and generate a response
The orchestrator combines the question with selected results in a prompt and sends it to the language model. The model generates a response, which the application returns to the user. A passage appearing in the prompt is evidence the model can use, not proof that its final answer is accurate or complete.
Rank #3
5. Evaluate and refine
Teams assess both the search results and the generated response, then adjust the data preparation, retrieval configuration, or prompt as needed. Evaluation should cover the full experience rather than treating a plausible-sounding response as evidence that the system works.
What should you evaluate in a RAG system?
Separate retrieval quality from response quality, then check whether the end-to-end result serves the user’s question. Microsoft names groundedness, completeness, utilization, and relevancy as possible response metrics; teams should select measures that fit their task and document configuration choices and evaluation results.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Retrieval: Does search return useful, relevant evidence for the question, including the right passages and enough context?
- Groundedness: Is the answer supported by the retrieved material, rather than making unsupported claims?
- Completeness and relevancy: Does the response address the question fully without drifting into unrelated information?
- Utilization: Does the model make appropriate use of the retrieved context?
- End-to-end behavior: Does the whole application provide a useful response, given its data pipeline, search, prompt, and model?
A weak answer can originate at different points. The collection may not contain the needed information; indexing or chunking may make it hard to find; search may return the wrong evidence; or the model may misread or fail to use relevant passages. Measuring only the final answer makes those causes difficult to distinguish.
How is agentic RAG different from standard RAG?
In standard RAG, orchestration follows a predetermined search-and-answer flow: receive a question, search a known source or index, assemble context, call the model, and return the response. This is often a natural fit when questions can be handled by searching a known collection.
Agentic RAG makes retrieval available as a tool an agent can choose to invoke. Depending on the design, the agent may select among sources, break a complex question into subquestions, or repeat searches as it works toward an answer. Microsoft suggests considering this approach when a fixed pipeline does not fit needs such as multistep reasoning or dynamic source selection.
| Design | How retrieval is used | What to evaluate |
|---|---|---|
| Standard RAG | A predetermined flow searches a configured source or index before the model generates an answer. | Retrieval and response quality, including groundedness, completeness, utilization, and relevancy. |
| Agentic RAG | An agent can decide when and how to use retrieval tools, potentially choosing sources or iterating through searches. | Response quality plus tool-selection accuracy, retrieval efficiency (including tool calls per request), and end-to-end latency by component. |
Agentic RAG is a design choice, not simply a more capable name for standard RAG. Giving an agent more control can address workflows that need flexible retrieval, but it also adds decisions about tool selection, efficiency, and latency. The added complexity is worthwhile only if the task calls for it.
Best Value
What should you consider when choosing a RAG approach?
- Data and search fit: Identify which formats and data structures must be searched, and test whether vector, full-text, hybrid, or sequential searches suit the actual questions.
- Operational control: Decide whether a managed service or a more customizable, self-managed architecture better matches your team’s needs. Google Cloud publishes examples covering managed vector search, database-backed vectors, and container-based architectures; these are implementation examples, not a neutral benchmark or universal recommendation. See Google Cloud Architecture Center: Generative AI with RAG.
- Quality and performance: Check whether retrieval finds useful evidence and answers remain relevant and grounded. For an agentic design, also assess tool selection and latency.
- Cost and governance: These are important deployment considerations, but there is no comparable current price basis here for recommending a vendor. Check the relevant provider’s current primary documentation before relying on prices, limits, regional availability, or security capabilities.
Microsoft, Google Cloud, and AWS all publish RAG architecture guidance. Their material can help explain particular implementation choices, but no cloud provider’s architecture should be treated as the universal standard. For another introductory overview, see Google Cloud: What is Retrieval-Augmented Generation (RAG)?
When is RAG useful—and what can it not do?
RAG is useful when an application needs to answer using information in an external collection, such as organization-specific documents, and can retrieve relevant passages from that source. It provides a way to bring that material into the model’s context when a question is asked.
It does not make missing or poor-quality data useful, ensure that search finds the right passages, or force the model to interpret evidence correctly. RAG should therefore be understood as a way to connect retrieval with generation—not as a guarantee that an answer is current, complete, or true. Its value depends on the task and on how the full system performs in evaluation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




