Build the chatbot as a retrieval-augmented generation (RAG) system: collect the documentation you allow it to use, split and index that content, retrieve the best passages for each question, and ask a language model to answer only from those passages. Return the answer with links to the original pages, refuse questions the documentation cannot support, and test the system before launch. This design is more dependable than asking a model to remember an entire website.
What the finished system does
A documentation chatbot has five cooperating parts:
- Content boundary: a defined set of public or authorized documentation pages, versions and sections.
- Ingestion: a crawler or publishing hook that extracts text, headings, code and metadata.
- Index: chunks are embedded and stored for semantic search, optionally alongside keyword search.
- Answer service: each question retrieves relevant chunks and supplies them to the response model as evidence.
- Interface and operations: a chat page, source links, authentication, rate limits, monitoring and an update process.
OpenAI’s Knowledge Retrieval blueprint summarizes the goal as: “Generate responses grounded in your data—with citations and evals for reliability.” Treat that as an architecture objective, not a guarantee that any particular model or prompt will be correct.
1. Define exactly what the bot may answer
Start with a content inventory instead of crawling the whole domain. Decide which documentation sections are authoritative and which pages must be excluded.
Recommended Free Tools
#1 Best Overall
Choose the corpus
- Include product guides, API references, tutorials, release notes and troubleshooting pages that are still supported.
- Exclude marketing copy, account pages, search results, comments, legal notices and obsolete versions unless users genuinely need them.
- Record the canonical URL, page title, product and version, section heading, publication or update time, and access policy for every page.
- Decide how tables, command examples, diagrams with alt text and embedded code should be represented. Preserve code as text; do not flatten it into navigation noise.
Handle versions and duplicates
Put version information in metadata and in the chunk text. A question about version 2.4 should not retrieve an unversioned page that describes version 1.9. Canonicalize URLs, remove tracking parameters, and deduplicate identical content. If two pages intentionally differ by platform, retain the platform label so citations explain the distinction.
Set an access policy
For private documentation, enforce the user’s authorization before retrieval. Do not put private chunks into a globally shared index and rely on the model to hide them. Public and private corpora can use separate collections or an authorization filter attached to every search.
2. Ingest pages and refresh the index
Indexing is a pipeline, not a one-time prompt. OpenAI’s retrieval documentation describes vector stores as indexes: files are chunked, embedded and indexed. In production, schedule ingestion or trigger it from your documentation build so changed and removed pages are reflected in the search layer.
Extraction workflow
- Fetch only approved URLs, respecting robots rules, authentication and crawl limits.
- Remove navigation, cookie notices, newsletter forms, chat widgets and repeated footers. Keep headings, paragraphs, lists, tables and code blocks.
- Normalize whitespace and retain the page title, URL, heading path, version, update timestamp and a stable document ID.
- Split text at logical boundaries. A heading and its explanation should remain together; do not split a command from the paragraph that explains its flags.
- Embed each chunk and upsert it with its metadata. Store a content hash so unchanged chunks do not need to be re-embedded.
- When a page changes, replace its old chunks. When a page is deleted or unpublished, remove its document ID from the index.
There is no universal chunk size, overlap, embedding model, top-k value or similarity threshold. Tune those values against your own questions, especially for long API references and pages containing many short code samples.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
3. A small, runnable Python RAG service
The following example demonstrates the complete retrieve-then-generate loop. It crawls same-host HTML pages from a seed URL, writes a JSON index containing embeddings and metadata, then answers questions with citations. It is intentionally simple: replace the JSON file with a managed vector store or Qdrant when you need concurrent writers, filtering, backups or a large corpus.
Install and configure
Use Python 3.10 or newer, then install openai, requests and beautifulsoup4. Set OPENAI_API_KEY and CHAT_MODEL to models available in your account. The embedding model can be changed with EMBEDDING_MODEL.
Save as chatbot.py
import argparse, json, math, os, re
from collections import deque
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
from openai import OpenAI
INDEX = 'docs-index.json'
client = OpenAI(api_key=os.environ['OPENAI_API_KEY'])
EMBEDDING_MODEL = os.getenv('EMBEDDING_MODEL', 'text-embedding-3-small')
CHAT_MODEL = os.environ['CHAT_MODEL']
def clean_page(url):
html = requests.get(url, timeout=30, headers={'User-Agent': 'DocsBot/1.0'}).text
soup = BeautifulSoup(html, 'html.parser')
for node in soup(['script', 'style', 'nav', 'footer', 'header', 'form']):
node.decompose()
title = soup.title.get_text(' ', strip=True) if soup.title else url
text = soup.get_text('n', strip=True)
text = re.sub(r'n{2,}', 'n', text)
return title, text, soup
def chunks(text, size=1400, overlap=200):
words = text.split()
out = []
step = max(1, size - overlap)
for start in range(0, len(words), step):
part = ' '.join(words[start:start + size])
if part:
out.append(part)
return out
def ingest(seed, limit=100):
host = urlparse(seed).netloc
queue, seen, records = deque([seed]), set(), []
while queue and len(seen) < limit:
url = queue.popleft().split('#')[0]
if url in seen or urlparse(url).netloc != host:
continue
seen.add(url)
try:
title, text, soup = clean_page(url)
except requests.RequestException:
continue
for part in chunks(text):
records.append({'url': url, 'title': title, 'text': part})
for link in soup.find_all('a', href=True):
target = urljoin(url, link['href']).split('#')[0]
if urlparse(target).netloc == host and target not in seen:
queue.append(target)
for record in records:
result = client.embeddings.create(model=EMBEDDING_MODEL, input=record['text'])
record['embedding'] = result.data[0].embedding
with open(INDEX, 'w', encoding='utf-8') as handle:
json.dump(records, handle)
print(f'Indexed {len(records)} chunks from {len(seen)} pages')
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
na = math.sqrt(sum(x * x for x in a))
nb = math.sqrt(sum(y * y for y in b))
return dot / (na * nb) if na and nb else 0.0
def answer(question, top_k=6):
with open(INDEX, encoding='utf-8') as handle:
records = json.load(handle)
query = client.embeddings.create(model=EMBEDDING_MODEL, input=question).data[0].embedding
ranked = sorted(records, key=lambda r: cosine(query, r['embedding']), reverse=True)[:top_k]
evidence = 'nn'.join(f"[{i}] {r['title']} — {r['url']}n{r['text']}" for i, r in enumerate(ranked, 1))
instructions = ('Answer only from the supplied documentation. If it does not contain the answer, '
'say that you could not find it and suggest a human support path. '
'Cite supporting sources as [number] after each material claim. Do not invent URLs.')
response = client.responses.create(model=CHAT_MODEL, instructions=instructions,
input=f'Question: {question}nnDocumentation:n{evidence}')
print(response.output_text)
parser = argparse.ArgumentParser()
parser.add_argument('command', choices=['ingest', 'ask'])
parser.add_argument('value')
args = parser.parse_args()
if args.command == 'ingest':
ingest(args.value)
else:
answer(args.value)
Run python chatbot.py ingest https://example.com/docs/, then python chatbot.py ask 'How do I rotate an API key?'. For a real site, replace the crawler with your documentation source of truth when possible; published Markdown or a build manifest is less noisy than scraping rendered pages.
4. Make citations trustworthy
Each retrieved record needs enough metadata to point back to the source page. At minimum, keep the canonical URL and title; section anchors, version and an excerpt make the citation more useful. The answer service should pass numbered evidence blocks to the model and render those numbers as links only after validating that they refer to retrieved records.
Tell the model to distinguish absence from uncertainty. If the retrieved passages do not answer the question, the correct response is a clear limitation, not a plausible-sounding guess. Provide a support link or contact route for that case. Citation correctness matters as much as answer correctness: a link to a page that merely mentions a feature is not evidence for a claim about its limits.
5. Select a retrieval stack
There is no single required database. Choose according to deployment constraints, data handling, existing operations knowledge and how much retrieval behavior you must customize.
| Option | What it provides | Best fit and trade-offs |
|---|---|---|
| Managed OpenAI retrieval | Vector stores, semantic search, file indexing and File Search guidance. | Fastest managed start; consider provider dependence, data handling and retrieval controls. The retrieval guide listed up to 1 GB across vector stores free and $0.10/GB/day beyond that when accessed in 2026; verify current pricing before purchase. |
| OpenAI Knowledge Retrieval starter kit | Config-first RAG workflow with File Search or local Qdrant, ChatKit and an evaluation harness. | Useful when you want citations and an evaluation flow quickly, with more customization than a purely hosted setup but more maintenance than a single API. |
| OpenSearch | A vector index, semantic retrieval and a conversational-agent tutorial. | Practical for teams already operating OpenSearch; plan for index operations and integration work. |
| Google Cloud GKE tutorial architecture | Files in Cloud Storage, document embeddings, a vector database and a semantic-search chatbot deployed on GKE. | Fits Google Cloud and Kubernetes expertise; operational complexity is higher than a small managed service. |
These sources demonstrate workflows rather than a head-to-head benchmark. Do not infer identical privacy, cost, latency or production readiness from the tutorials.
6. Build the chat interface safely
Put the model call behind your own server endpoint. Browser code should send the user’s message to your backend; it should never contain a provider secret. The endpoint should authenticate users when documentation is restricted, apply per-user or per-IP rate limits, enforce maximum question and context sizes, and log request IDs without storing secrets.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Return a structured response such as {answer, citations, refusal, request_id}. The interface should show a loading state, recoverable error messages and clickable source titles. Make the widget keyboard accessible, expose status changes to assistive technology, and keep citations visible on mobile. Sanitize rendered Markdown or HTML so documentation text cannot inject scripts into the chat page.
7. Evaluate before deployment
Create a test set from real support questions and documentation search logs. Include exact product and version questions, questions requiring two pages, ambiguous wording, code questions, unsupported questions and attempts to make the model ignore its evidence.
- Answer correctness: does the response match the documentation?
- Retrieval quality: did the returned chunks contain the needed passage?
- Citation correctness: does every cited page support the nearby claim?
- Abstention: does the bot say it cannot answer when the corpus is silent?
- Latency and cost: measure retrieval, generation and total response time separately.
Run the set before launch and whenever documentation, prompts, models, chunking or retrieval settings change. The starter kit’s evaluation workflow is a useful model, but your pass criteria must reflect your product’s risk. A bot answering billing or security questions needs stricter review than one answering a harmless tutorial.
8. Reliability, performance and cost controls
Reduce latency
Embed documents during ingestion, not during a user request. Cache repeated question embeddings and retrieval results for content that changes infrequently. Keep retrieved context to the passages needed for the answer, and stream generated text only after your service has completed authorization and retrieval.
Keep answers fresh
Store a source revision or content hash with every chunk. Run a scheduled refresh and trigger an immediate refresh from the documentation build. Monitor for pages that disappeared, changed versions or repeatedly produce no usable chunks.
Control spend
Embedding costs occur mainly during ingestion and updates; generation costs occur per conversation. Avoid re-embedding unchanged content, cap crawl scope, limit maximum retrieved chunks and record token usage. Storage pricing and model prices change, so calculate using the current provider documentation rather than a fixed estimate in code.
Protect against prompt injection
Documentation can contain text that looks like an instruction. Mark retrieved material as untrusted evidence, keep system instructions outside it, and test pages containing requests to reveal secrets or ignore policy. Never allow documentation text to select tools, change authorization or override the refusal rule.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.9. Troubleshooting common failures
The bot invents an answer
Check whether retrieval returned relevant chunks. Improve page cleaning, metadata filters and chunk boundaries; then strengthen the instruction to answer only from evidence and to abstain when evidence is missing. Do not solve this solely by lowering the model temperature.
Citations point to the wrong page
Ensure each evidence block has a stable number for the duration of the request. Preserve canonical URLs and validate citation numbers server-side before rendering links. Duplicate pages and redirects should be collapsed during ingestion.
New documentation is ignored
Inspect the refresh job, content hashes and deletion handling. A successful crawler run is not proof that the index changed; report the number of added, replaced and removed chunks and alert when it is unexpectedly zero.
Answers are slow or exceed context limits
Measure retrieval and generation separately. Reduce duplicate chunks, cap top_k, shorten boilerplate and use section-aware chunking. If a question truly spans many pages, retrieve in stages and summarize evidence before the final response.
Private content leaks
Verify authorization before search and test with two accounts that have different permissions. Filter by tenant, product and version in the data layer, not in the final prompt. Review logs for accidental storage of document text or credentials.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsOr skip the browser setup
If you need screenshots of documentation pages for visual QA, release notes or a chatbot demo, ScreenshotNeo can capture a page with one request. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for all options. cURL:
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/docs/ -o docs.webp
Python:
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://example.com/docs/'}, timeout=90)
open('docs.webp', 'wb').write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/docs/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account to try it.
10. A launch checklist
- Corpus scope, versions and exclusions are documented.
- Every chunk has a canonical URL, title and access metadata.
- Changed and deleted pages are reflected in the index.
- Answers cite retrieved evidence and abstain when evidence is insufficient.
- Secrets stay server-side; authentication, rate limits and output sanitization are enabled.
- Evaluation covers correctness, retrieval, citations, refusal and latency.
- Monitoring records stale pages, failed retrievals, user feedback and token or storage costs.
Frequently Asked Questions
Can a documentation chatbot answer questions about several product versions?
Yes, if version is stored as retrieval metadata and the user’s requested version is a mandatory filter. If no version is specified, ask a clarifying question or state which version the cited page covers.
How should I remove a page permanently?
Delete every chunk associated with its stable document ID, purge cached retrieval results, and add a test that searches for a distinctive sentence from the removed page.
What should the bot do when a user asks for account-specific data?
Do not retrieve or expose that data from a general documentation index. Route the user to an authenticated support or account workflow that can enforce identity and authorization.
Is semantic search enough for exact command names?
Use semantic retrieval together with exact or keyword matching when commands, error codes and parameter names matter. Merge and rerank both result sets before generation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

