October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk7 min

How to Build a RAG System from Scratch in Python: Chunk, Embed, Retrieve, Cite

Build an inspectable Python RAG pipeline that keeps document provenance through chunking, embedding, semantic retrieval, answer generation, and citation rendering.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A minimal retrieval-augmented generation (RAG) system has four jobs: split source documents into traceable chunks, embed and store those chunks, retrieve relevant passages for a question, and generate an answer whose citations map back to the original sources. The example below keeps those steps visible with Python, an in-memory vector index, and an embedding API; it is a teaching baseline, not a production-ready service.

What a RAG pipeline does

RAG gives a language model selected source text at answer time instead of asking it to rely only on information encoded during training. The pipeline has two paths:

As an Amazon Associate I earn from qualifying purchases.

  • Indexing: parse documents, divide them into chunks, embed each chunk, and store its text, vector, and provenance.
  • Answering: embed a question, find relevant stored chunks, pass them to a language model, and render citations from the chunks actually retrieved.

Embeddings are numeric representations of text. Similarity search uses them to find passages that are semantically related to a query; it does not establish that a passage is true or that it answers the question. The OpenAI embeddings guide demonstrates creating embeddings in Python and describes saving them in a vector database. Its retrieval guide shows search over a managed vector store. These are OpenAI examples, not requirements for every RAG system.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a local or managed retrieval path

The code here makes chunking, similarity ranking, and source mapping inspectable. A managed retrieval service can take over some indexing and search work, but introduces provider-specific APIs and may expose fewer details of its internal pipeline. Neither approach is inherently more accurate: compare them on the same representative questions, including whether each answer cites the right passage.

Decision Local implementation Managed retrieval
Control and inspection You can directly inspect parsing, chunk boundaries, vector math, and citation mapping. The service automates more retrieval work; some implementation details may be hidden.
Operations You choose and maintain storage, indexing, updates, and scaling. The provider supplies hosted retrieval components, but you depend on its interfaces and data-handling terms.
Evaluation Test relevance and citation correctness with a fixed question set. Use the same questions and correctness checks; no cross-provider quality result follows from using a managed service.
Cost and scale Measure your storage and compute needs at realistic corpus and query volumes. Check current provider pricing and test realistic volumes; there is no universal cost comparison.

For small examples, an in-memory matrix and cosine similarity keep the mechanics clear. A persistent vector index or hosted store becomes useful when you need durable storage, metadata filtering, updates, or scale. Those choices should follow the needs of the corpus and workload, not an assumed performance advantage.

Build the data path around traceable records

Keep source metadata with every chunk

Before parsing, assign each document a stable ID and record its locator, such as a filename or URL, and its title if available. A chunk should carry its own stable ID, text, document ID, and useful location data: for example, a heading, page number, or character offsets. Retain headings and table labels when they are needed to interpret a passage. Keep the original text or a dependable reference to it alongside the vector.

This record design is what makes a search result usable as a citation. A vector match by itself cannot tell a reader which source passage supports an answer. OpenAI’s vector-store file reference documents file metadata and chunking options; the fields available in a particular managed workflow should be checked rather than assumed to include every provenance detail your application needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse and normalize without losing meaning

Use a parser appropriate to each input format and report files that fail to parse. Normalize whitespace, but preserve meaningful structure and capture source locations before chunking where possible. If a PDF parser extracts text without reliable page or section information, do not claim citations to page-level locations unless you can verify that mapping.

How to chunk documents for RAG

Start at natural boundaries such as headings and paragraphs, then split sections that exceed a chosen token or character ceiling. A single universal chunk size is not established: the right setting depends on the document structure, the kinds of questions, and how much surrounding context an answer needs.

  • Chunks that are too large can mix unrelated subjects and make a focused match less useful.
  • Chunks that are too small can separate a fact from the heading, definition, or qualification needed to understand it.
  • Overlap can preserve context across split boundaries, but also duplicates stored text and may return redundant context.

Keep chunk IDs and source offsets when splitting. Evaluate a few plausible strategies against questions whose answers have known passages; inspect both the retrieved text and its citation location before choosing settings.

For comparison, OpenAI-managed vector-store files document an automatic chunking strategy with a maximum of 800 tokens and 400-token overlap. Its static strategy allows a maximum chunk size from 100 to 4,096 tokens, and overlap cannot exceed half the maximum chunk size. These are provider-specific defaults and constraints, not a general recommendation for local chunking. See the vector-store file reference for the current API details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Embed and store chunks

An embedding API accepts text and returns a vector. The following is the documented OpenAI Python call pattern; it assumes the SDK is installed and the client is configured with credentials. Check the current SDK and model documentation before using it in an application.

from openai import OpenAI

client = OpenAI()
response = client.embeddings.create(
    input="Text to embed",
    model="text-embedding-3-small",
)
vector = response.data[0].embedding

Embed each chunk and persist a record containing at least its chunk ID, text or text reference, source metadata, embedding model, and vector. Use the same embedding model for indexed chunks and query vectors. If you change models, rebuild or otherwise consistently migrate the index; vectors from different models should not be treated as directly comparable.

OpenAI’s embeddings guide lists default dimensions of 1,536 for text-embedding-3-small and 3,072 for text-embedding-3-large, and an 8,192-token maximum input for the models listed there. These are product specifications that can change, not universal embedding limits. Consult the current embeddings guide before setting application limits.

Retrieve relevant context for a question

Embed the user’s question with the same model used for the chunks. Compare that vector with the stored vectors, rank the results, and inspect the top candidates before choosing the context that will fit into the generation prompt. The embeddings guide recommends cosine similarity and notes that OpenAI embeddings are unit-normalized; a vector database can perform the search instead. The retrieval guide demonstrates a natural-language query against a managed vector store.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import numpy as np

def cosine_similarity(a, b):
    a = np.asarray(a, dtype=float)
    b = np.asarray(b, dtype=float)
    denominator = np.linalg.norm(a) * np.linalg.norm(b)
    if denominator == 0:
        return 0.0
    return float(np.dot(a, b) / denominator)

query_vector = embed(question)
ranked = sorted(
    chunks,
    key=lambda chunk: cosine_similarity(query_vector, chunk["embedding"]),
    reverse=True,
)
context_chunks = ranked[:4]

Here, embed represents the embedding call and chunks records containing vectors and provenance. The value four is only an illustrative selection count: choose the number and total context size according to the model’s prompt limits and evaluation results. Retrieving more candidates than you ultimately include can help you inspect ranking and relevance. A high similarity score is not proof that a passage supports an answer. Keyword or hybrid retrieval can be considered when exact names, IDs, dates, or rare terms matter, but its ranking behavior must be evaluated for your corpus.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Generate an answer grounded in retrieved passages

Pass the question and selected passages to a generation model as structured context. Instruct it to answer from those passages, say when the evidence is insufficient, and associate factual claims with source identifiers. Keep each passage’s metadata in application data rather than flattening it irreversibly into a prompt.

context = [
    {
        "chunk_id": chunk["chunk_id"],
        "text": chunk["text"],
        "source": chunk["source"],
        "location": chunk["location"],
    }
    for chunk in context_chunks
]

# Send the question and context to your generation model.
# Require claims to identify supporting chunk_id values.
# Handle an answer with insufficient supporting evidence.

This is application logic, not a universal prompt recipe. The retrieval and embeddings APIs provide useful primitives; your code must decide how to format evidence, how to handle conflicting passages, and when not to answer.

How to cite sources in an AI answer

Render citations by mapping each cited chunk ID back to the retrieved record and then to its source locator and location. Show citations next to the claims they support. Reject any identifier that was not among the retrieved chunks, and provide a clear fallback when no passage supports an answer. If a chunk points to a file, link to the file or show its name and location; if it points to a web document, retain the source URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Collect the IDs of chunks actually supplied to the generation model.
  2. Validate every ID returned for citation against that set; do not display unknown IDs.
  3. Resolve each valid ID to its source URL or filename and stored location.
  4. Render the citation beside the supported claim and verify that it opens or identifies the intended passage.
  5. If the model’s claim has no supporting retrieved passage, omit it or state that the available sources do not establish it.

OpenAI’s file-search guide documents file citations in generated responses. A custom pipeline needs its own equivalent source mapping and validation; generating a citation-looking label is not enough to make it accurate.

Test the pipeline before trusting its answers

Build a small evaluation set with questions whose supporting source passages are known. For each question, inspect whether retrieval finds those passages, whether the answer stays within their evidence, and whether every citation resolves to the correct source and location. Include questions with no answer in the corpus to check the insufficient-evidence path. Record errors such as parse failures, missing metadata, irrelevant retrievals, and invalid citation IDs so you can distinguish indexing problems from generation problems.

There are no comparative quality, latency, or cost benchmarks established here for local versus hosted retrieval. Measure those dimensions against your own representative corpus, query volume, and operational constraints before selecting an architecture.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.