Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a chatbot that must answer questions about changing website content, “training” usually means building a permission-aware web crawl and a searchable retrieval index—not changing a model’s weights. Crawl pages you are entitled to use, turn them into clean passages, retrieve relevant passages for each question, and require the chatbot to ground its answer in that context. Then evaluate the whole pipeline before launch.

What “training on scraped content” means

A language model’s built-in knowledge is not a live copy of your website. If a product page, policy, or help article changes, fine-tuning a model does not automatically update that fact. A more maintainable approach is retrieval-augmented generation (RAG): the model receives relevant passages from a separate, refreshable knowledge base when it answers.

The pipeline has two distinct paths. During ingestion, you crawl allowed pages, extract and clean their content, split it into passages, and index those passages. At question time, you search the index, pass the best matches and their source details to the model, and ask it to answer from that evidence. OpenAI’s Retrieval documentation describes semantic search over vector stores, which can surface relevant passages even when a question uses few of the same words as the source.

Fine-tuning is a different tool. It may help with consistent tone, formatting, or task behavior when evaluation shows that the model’s response behavior is the problem. It does not give you a refreshable index of current web facts. OpenAI’s optimization guidance recommends choosing techniques based on the observed failure; its fine-tuning documentation has described the platform as winding down and unavailable to new users. Availability can change, so check the provider’s current documentation before planning around fine-tuning.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define what the chatbot is allowed to know

Before writing a crawler, decide what the chatbot should answer and which sources it may use. Write down:

  • The permitted domains and URL paths, such as your own help center but not customer forums.
  • Content types and languages to include, and pages or personal information to exclude.
  • The questions users are expected to ask and how current answers need to be.
  • How often each source should be refreshed, and how deleted or changed pages will be handled.
  • Who can access each indexed document, if the site contains public and restricted material.

Prefer a content export, official API, feed, sitemap, or explicit license when one is available and suitable. Scrape only when the source and applicable rules allow your intended use. Publicly reachable does not automatically mean free to store indefinitely, republish, or use for any purpose. Check the site’s terms, applicable licenses and laws, and privacy implications. Read and honor crawler instructions such as robots.txt as a baseline, but do not treat them as complete legal authorization.

Keep a source manifest with each page’s canonical URL, retrieval time, response status, and any relevant access or license notes. It gives you a traceable way to investigate stale answers and remove material if needed.

Crawl a bounded set of pages politely

Start with an explicit allowlist and a small page limit. Identify your crawler with a clear user agent, set a delay between requests, avoid unnecessary concurrency, and slow down or stop if the server begins returning errors. Scrapy’s AutoThrottle documentation describes adjusting per-site delays based on response latency; the aim is to be more considerate than a zero-delay default. Whatever crawler you use, set limits rather than assuming the defaults suit the site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The following small Python example fetches only the seed URLs you explicitly list. It checks robots.txt for those URLs, waits between requests, and writes extracted text with source metadata to JSON Lines. It is a starting point, not a full crawler: it does not discover links, handle JavaScript-rendered content, or replace a review of the site’s terms and applicable rules.

from html.parser import HTMLParser
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
from urllib.request import Request, urlopen
import json
import time

USER_AGENT = "ExampleKnowledgeBot/1.0 (contact: [email protected])"
ALLOWED_HOSTS = {"example.com", "www.example.com"}
SEED_URLS = [
    "https://example.com/help/",
    "https://example.com/help/shipping",
]
DELAY_SECONDS = 2
MAX_PAGES = 20

class TextExtractor(HTMLParser):
    BLOCKED = {"script", "style", "noscript", "svg"}
    def __init__(self):
        super().__init__()
        self.skip = 0
        self.parts = []
    def handle_starttag(self, tag, attrs):
        if tag in self.BLOCKED:
            self.skip += 1
    def handle_endtag(self, tag):
        if tag in self.BLOCKED and self.skip:
            self.skip -= 1
    def handle_data(self, data):
        if not self.skip and data.strip():
            self.parts.append(" ".join(data.split()))

def allowed_by_robots(url):
    parts = urlparse(url)
    robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
    parser = RobotFileParser(robots_url)
    try:
        parser.read()
    except Exception as exc:
        raise RuntimeError(f"Could not read {robots_url}; review before crawling") from exc
    return parser.can_fetch(USER_AGENT, url)

seen = set()
with open("pages.jsonl", "w", encoding="utf-8") as output:
    for url in SEED_URLS[:MAX_PAGES]:
        host = urlparse(url).hostname
        if host not in ALLOWED_HOSTS or url in seen:
            continue
        seen.add(url)
        if not allowed_by_robots(url):
            print(f"Skipped by robots.txt: {url}")
            continue
        request = Request(url, headers={"User-Agent": USER_AGENT})
        try:
            with urlopen(request, timeout=20) as response:
                status = response.status
                content_type = response.headers.get("Content-Type", "")
                if status != 200 or "text/html" not in content_type:
                    print(f"Skipped status/content type {status}: {url}")
                    continue
                html = response.read().decode("utf-8", errors="replace")
            extractor = TextExtractor()
            extractor.feed(html)
            record = {
                "url": url,
                "retrieved_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
                "status": status,
                "text": " ".join(extractor.parts),
            }
            output.write(json.dumps(record, ensure_ascii=False) + "n")
            print(f"Saved: {url}")
        except Exception as exc:
            print(f"Failed {url}: {exc}")
        time.sleep(DELAY_SECONDS)

Run it with Python 3: save the script as crawl.py, put your allowed seed URLs and hostnames in the lists, then run python crawl.py. Inspect pages.jsonl before indexing it. If robots.txt cannot be read, this example stops rather than silently treating the failure as permission to proceed. The extractor is intentionally basic: production extraction should preserve meaningful headings, tables, and lists, and should remove repeated navigation and consent text without discarding useful content.

Or skip the browser setup

For a visual capture of a rendered page, ScreenshotNeo can return a screenshot from one GET request. A screenshot is an image, not an extracted text document for RAG; use an HTML/content extraction path for a text knowledge base, or add your own OCR step if your use case genuinely needs text from pixels.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/help/ -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/help/"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/help/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are not billed; its MCP server gives AI agents screenshot tools; and the free plan includes 1,000 screenshots per month with no card, while paid plans start at $5 for 3,000. Learn about ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Clean, divide, and index the content

Raw page text is rarely ready for retrieval. Normalize encoding and whitespace, identify the language, remove exact and near duplicates, and retain meaningful structure. Exclude sensitive personal information that is not necessary for the chatbot’s purpose. Set retention and deletion procedures for source files and derived index entries; cleaning alone does not make a dataset compliant or safe.

Divide pages into passages that make sense on their own. Splitting by headings or other document structure is usually more useful than cutting at arbitrary character boundaries. Include enough context for a passage to identify its subject, but avoid huge chunks that bury the relevant sentence. Store metadata with every passage: source URL, page title, heading, crawl time, language, and access classification are useful starting points.

Embed or otherwise index those passages in a retrieval system. OpenAI’s Retrieval guide describes vector stores as indices and semantic search as a way to find conceptually related passages even when keyword overlap is low. Chunking and ranking affect results, so tune them against your own questions instead of assuming one default works for every site.

Retrieve evidence and generate a grounded answer

For each user question, retrieve a small set of relevant passages and pass their text and source metadata to the model. If exact product names, policy terms, or identifiers matter, consider combining semantic search with keyword retrieval. Tell the model to answer only when the retrieved context supports an answer, distinguish source-backed statements from inference, and ask for clarification or abstain when evidence is inadequate. Where appropriate, return links to the source pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieved web pages are evidence, not instructions to the chatbot. A page might contain text that attempts to override system behavior or request secrets. Treat crawled content as untrusted input: the model should use it as material to answer the user, not obey commands embedded in it.

Evaluate before launch and after every meaningful change

Build a representative question set before tuning the system. Include straightforward factual questions, paraphrases, questions about updated pages, conflicting source pages, questions whose answer is absent, and pages containing instruction-like text. Check retrieval and generation separately:

  • Retrieval: Did the index return the page and passage that actually answer the question?
  • Answer: Is every factual claim supported by the retrieved material, and does the answer handle missing or conflicting evidence appropriately?
  • Citations: Do returned links point to passages that support the claims beside them?
  • Freshness: After a page changes or disappears, does the chatbot stop using stale content?

OpenAI’s knowledge-retrieval workflow places evaluations before deployment. Repeat your own evaluations after changing extraction, chunking, ranking, prompts, or the model; each can alter the result even when the crawl is unchanged.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Refresh the index and manage its data

Choose a refresh schedule that matches how quickly sources change and the load the site can reasonably tolerate. Compare new page versions with indexed versions, update changed passages, and expire removed pages. Deletion must propagate through derived chunks and embeddings, not just the original HTML file. Keep a record of crawl decisions and source URLs so you can trace outdated or disputed material.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep user conversations separate from the scraped corpus unless there is a clear, disclosed, lawful reason to combine them. Provider data controls apply to the data sent to that provider; they do not settle your own storage, application logs, retention, access-control, or legal responsibilities. OpenAI’s platform documentation describes provider-specific API data controls and default abuse-monitoring retention, with additional controls for eligible customers. Check the current terms for your provider and account rather than generalizing one vendor’s policy to another.

RAG or fine-tuning: choose based on the failure

Decision Retrieval over a scraped knowledge base Fine-tuning
Main purpose Supply current or external facts at answer time Change response behavior, style, format, or task performance
Updating source facts Refresh and re-index the relevant documents Requires another training process and does not automatically refresh facts
Traceability Can return retrieved passages and URLs Model weights alone do not show which source produced an answer
Typical tuning work Improve ingestion, chunking, retrieval, ranking, prompts, and evaluation Curate examples, train, validate, and check for regressions
Best next step Use when the issue is missing or stale reference context Consider only when evaluations show a behavior problem examples can improve

These are engineering distinctions, not guarantees: results depend on the model, data, evaluation, and deployment. Start with the simplest approach that addresses the observed failure rather than treating scraping and fine-tuning as mandatory consecutive steps.

Common problems and fixes

  • The chatbot misses content visible in a browser. The page may be rendered by JavaScript, or the text extractor may have discarded important structure. Confirm whether the allowed source offers an export or API; otherwise use a rendering-capable acquisition method, then test the extracted text before indexing it.
  • Answers use old information. Check crawl timestamps, refresh cadence, duplicate versions, and whether changed or removed content is propagated through the index. A correct retrieval system cannot return an update it has not ingested.
  • Search returns plausible but wrong passages. Review chunk boundaries and metadata, then tune retrieval and ranking against a labeled question set. Consider keyword matching alongside semantic search for exact names and codes.
  • The model invents details despite retrieved context. Tighten the instruction to answer only from evidence, reduce irrelevant retrieved passages, and test abstention behavior on questions with no answer in the index.
  • The source returns errors or crawling gets blocked. Verify the scope, user agent, terms, and robots instructions; reduce request rate and concurrency, and stop if errors rise. Do not attempt to bypass access controls.
  • Indexing duplicates or removed pages persist. Canonicalize URLs, track document versions, deduplicate content, and implement deletion through every derived representation.

Privacy and crawler controls are separate questions

OpenAI describes its own foundation-model development as using publicly available internet content, partner information, and material from human trainers and researchers, alongside filtering efforts; that description is about OpenAI’s practices, not a legal rule or authorization for your crawler. Its crawler documentation also distinguishes OAI-SearchBot, used for discovery for ChatGPT search, from GPTBot, which may crawl pages for possible foundation-model training. Those controls have different purposes; allowing one should not be treated as allowing the other. OpenAI notes that changes to robots.txt settings can take about 24 hours to affect search behavior.

Frequently Asked Questions

Can I use a screenshot as the chatbot’s source document?

A screenshot contains pixels rather than the structured page text a typical RAG index searches. Extract text from the page itself where possible; OCR is an additional step if image-derived text is required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I let the chatbot learn from user conversations automatically?

Not by default. Keeping conversations separate avoids silently mixing user data into the source corpus; combining them needs a clear purpose, disclosure, and appropriate legal and retention controls.

What should the chatbot do when two indexed pages disagree?

Make the conflict visible rather than silently choosing a claim. Preserve source and update metadata so your product can apply an explicit precedence policy or ask a human to resolve authoritative-source questions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.