A minimal retrieval-augmented generation (RAG) system has four jobs: split source documents into traceable chunks, embed and store those chunks, retrieve relevant passages for a question, and generate an answer whose citations map back to the original sources. The example below keeps those steps visible with Python, an in-memory vector index, and an embedding API; it is a teaching baseline, not a production-ready service.
What a RAG pipeline does
RAG gives a language model selected source text at answer time instead of asking it to rely only on information encoded during training. The pipeline has two paths:
As an Amazon Associate I earn from qualifying purchases.
- Indexing: parse documents, divide them into chunks, embed each chunk, and store its text, vector, and provenance.
- Answering: embed a question, find relevant stored chunks, pass them to a language model, and render citations from the chunks actually retrieved.
Embeddings are numeric representations of text. Similarity search uses them to find passages that are semantically related to a query; it does not establish that a passage is true or that it answers the question. The OpenAI embeddings guide demonstrates creating embeddings in Python and describes saving them in a vector database. Its retrieval guide shows search over a managed vector store. These are OpenAI examples, not requirements for every RAG system.
Free tools Windows power users keep installed
One-click scans. No signup required.
Choose a local or managed retrieval path
The code here makes chunking, similarity ranking, and source mapping inspectable. A managed retrieval service can take over some indexing and search work, but introduces provider-specific APIs and may expose fewer details of its internal pipeline. Neither approach is inherently more accurate: compare them on the same representative questions, including whether each answer cites the right passage.
#1 Best Overall
| Decision | Local implementation | Managed retrieval |
|---|---|---|
| Control and inspection | You can directly inspect parsing, chunk boundaries, vector math, and citation mapping. | The service automates more retrieval work; some implementation details may be hidden. |
| Operations | You choose and maintain storage, indexing, updates, and scaling. | The provider supplies hosted retrieval components, but you depend on its interfaces and data-handling terms. |
| Evaluation | Test relevance and citation correctness with a fixed question set. | Use the same questions and correctness checks; no cross-provider quality result follows from using a managed service. |
| Cost and scale | Measure your storage and compute needs at realistic corpus and query volumes. | Check current provider pricing and test realistic volumes; there is no universal cost comparison. |
For small examples, an in-memory matrix and cosine similarity keep the mechanics clear. A persistent vector index or hosted store becomes useful when you need durable storage, metadata filtering, updates, or scale. Those choices should follow the needs of the corpus and workload, not an assumed performance advantage.
Build the data path around traceable records
Keep source metadata with every chunk
Before parsing, assign each document a stable ID and record its locator, such as a filename or URL, and its title if available. A chunk should carry its own stable ID, text, document ID, and useful location data: for example, a heading, page number, or character offsets. Retain headings and table labels when they are needed to interpret a passage. Keep the original text or a dependable reference to it alongside the vector.
This record design is what makes a search result usable as a citation. A vector match by itself cannot tell a reader which source passage supports an answer. OpenAI’s vector-store file reference documents file metadata and chunking options; the fields available in a particular managed workflow should be checked rather than assumed to include every provenance detail your application needs.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #2
Parse and normalize without losing meaning
Use a parser appropriate to each input format and report files that fail to parse. Normalize whitespace, but preserve meaningful structure and capture source locations before chunking where possible. If a PDF parser extracts text without reliable page or section information, do not claim citations to page-level locations unless you can verify that mapping.
How to chunk documents for RAG
Start at natural boundaries such as headings and paragraphs, then split sections that exceed a chosen token or character ceiling. A single universal chunk size is not established: the right setting depends on the document structure, the kinds of questions, and how much surrounding context an answer needs.
- Chunks that are too large can mix unrelated subjects and make a focused match less useful.
- Chunks that are too small can separate a fact from the heading, definition, or qualification needed to understand it.
- Overlap can preserve context across split boundaries, but also duplicates stored text and may return redundant context.
Keep chunk IDs and source offsets when splitting. Evaluate a few plausible strategies against questions whose answers have known passages; inspect both the retrieved text and its citation location before choosing settings.
For comparison, OpenAI-managed vector-store files document an automatic chunking strategy with a maximum of 800 tokens and 400-token overlap. Its static strategy allows a maximum chunk size from 100 to 4,096 tokens, and overlap cannot exceed half the maximum chunk size. These are provider-specific defaults and constraints, not a general recommendation for local chunking. See the vector-store file reference for the current API details.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Embed and store chunks
An embedding API accepts text and returns a vector. The following is the documented OpenAI Python call pattern; it assumes the SDK is installed and the client is configured with credentials. Check the current SDK and model documentation before using it in an application.
from openai import OpenAI
client = OpenAI()
response = client.embeddings.create(
input="Text to embed",
model="text-embedding-3-small",
)
vector = response.data[0].embedding
Embed each chunk and persist a record containing at least its chunk ID, text or text reference, source metadata, embedding model, and vector. Use the same embedding model for indexed chunks and query vectors. If you change models, rebuild or otherwise consistently migrate the index; vectors from different models should not be treated as directly comparable.
OpenAI’s embeddings guide lists default dimensions of 1,536 for text-embedding-3-small and 3,072 for text-embedding-3-large, and an 8,192-token maximum input for the models listed there. These are product specifications that can change, not universal embedding limits. Consult the current embeddings guide before setting application limits.
Retrieve relevant context for a question
Embed the user’s question with the same model used for the chunks. Compare that vector with the stored vectors, rank the results, and inspect the top candidates before choosing the context that will fit into the generation prompt. The embeddings guide recommends cosine similarity and notes that OpenAI embeddings are unit-normalized; a vector database can perform the search instead. The retrieval guide demonstrates a natural-language query against a managed vector store.
import numpy as np
def cosine_similarity(a, b):
a = np.asarray(a, dtype=float)
b = np.asarray(b, dtype=float)
denominator = np.linalg.norm(a) * np.linalg.norm(b)
if denominator == 0:
return 0.0
return float(np.dot(a, b) / denominator)
query_vector = embed(question)
ranked = sorted(
chunks,
key=lambda chunk: cosine_similarity(query_vector, chunk["embedding"]),
reverse=True,
)
context_chunks = ranked[:4]
Here, embed represents the embedding call and chunks records containing vectors and provenance. The value four is only an illustrative selection count: choose the number and total context size according to the model’s prompt limits and evaluation results. Retrieving more candidates than you ultimately include can help you inspect ranking and relevance. A high similarity score is not proof that a passage supports an answer. Keyword or hybrid retrieval can be considered when exact names, IDs, dates, or rare terms matter, but its ranking behavior must be evaluated for your corpus.
Best Value
Generate an answer grounded in retrieved passages
Pass the question and selected passages to a generation model as structured context. Instruct it to answer from those passages, say when the evidence is insufficient, and associate factual claims with source identifiers. Keep each passage’s metadata in application data rather than flattening it irreversibly into a prompt.
context = [
{
"chunk_id": chunk["chunk_id"],
"text": chunk["text"],
"source": chunk["source"],
"location": chunk["location"],
}
for chunk in context_chunks
]
# Send the question and context to your generation model.
# Require claims to identify supporting chunk_id values.
# Handle an answer with insufficient supporting evidence.
This is application logic, not a universal prompt recipe. The retrieval and embeddings APIs provide useful primitives; your code must decide how to format evidence, how to handle conflicting passages, and when not to answer.
How to cite sources in an AI answer
Render citations by mapping each cited chunk ID back to the retrieved record and then to its source locator and location. Show citations next to the claims they support. Reject any identifier that was not among the retrieved chunks, and provide a clear fallback when no passage supports an answer. If a chunk points to a file, link to the file or show its name and location; if it points to a web document, retain the source URL.
- Collect the IDs of chunks actually supplied to the generation model.
- Validate every ID returned for citation against that set; do not display unknown IDs.
- Resolve each valid ID to its source URL or filename and stored location.
- Render the citation beside the supported claim and verify that it opens or identifies the intended passage.
- If the model’s claim has no supporting retrieved passage, omit it or state that the available sources do not establish it.
OpenAI’s file-search guide documents file citations in generated responses. A custom pipeline needs its own equivalent source mapping and validation; generating a citation-looking label is not enough to make it accurate.
Test the pipeline before trusting its answers
Build a small evaluation set with questions whose supporting source passages are known. For each question, inspect whether retrieval finds those passages, whether the answer stays within their evidence, and whether every citation resolves to the correct source and location. Include questions with no answer in the corpus to check the insufficient-evidence path. Record errors such as parse failures, missing metadata, irrelevant retrievals, and invalid citation IDs so you can distinguish indexing problems from generation problems.
There are no comparative quality, latency, or cost benchmarks established here for local versus hosted retrieval. Measure those dimensions against your own representative corpus, query volume, and operational constraints before selecting an architecture.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




