Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use LangChain to turn web pages into searchable Document objects, index those documents for retrieval-augmented generation (RAG), and—only when the workflow needs dynamic decisions—let an agent choose retrieval or other tools. A dependable implementation keeps ingestion, indexing, question-time retrieval, and agent orchestration as separate layers.

What LangChain contributes to a web-knowledge workflow

LangChain treats retrieval as runtime access to information that a language model does not reliably contain, including private, recent, or application-specific data. Its retrieval components are modular: document loaders ingest sources, text splitters create searchable chunks, embedding models convert chunks to vectors, vector stores persist those vectors, and retrievers return relevant chunks for a question.

That modularity lets you replace a loader, splitter, embedding provider, or vector store without redesigning the whole application. It does not make every website scrapeable with one universal connector. A loader is an ingestion interface; the extraction mechanism and dependencies depend on the source integration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose two-step or agentic RAG first

The most important design decision is whether retrieval is always required or whether the model should decide when to retrieve and which tool to use.

Decision axis Two-step RAG Agentic RAG
Retrieval timing Always before generation The agent chooses when and how to retrieve
Control Higher Lower
Flexibility Lower Higher
Latency profile Generally more predictable Variable
Good fit FAQs and documentation bots Research assistants using multiple tools

Use two-step RAG when every question should search the same knowledge base. It is easier to test, cache, and operate. Use agentic RAG when a task may require several tools, conditional searches, calculations, or a decision about whether retrieval is needed. The additional flexibility brings a less predictable number of model and tool calls. A hybrid can add validation after an agent gathers evidence.

Ingest web pages as documents

Python loader example

Install the packages required by your selected integrations. Package names and import paths evolve, so pin and verify versions in your project before deployment.

pip install -U langchain langchain-community langchain-openai langchain-chroma beautifulsoup4

This example uses the community web loader, then prints the extracted text and metadata before indexing it:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from langchain_community.document_loaders import WebBaseLoader

urls = [
    "https://example.com/guide",
    "https://example.com/reference",
]
loader = WebBaseLoader(urls)
documents = loader.load()

for document in documents:
    print(document.metadata)
    print(document.page_content[:500])

Inspecting the result matters. If the page is mostly navigation, an interstitial, or an empty shell rendered by JavaScript, fix extraction before embedding. The loader’s behavior is source-specific; do not assume the same code extracts every site.

Official JavaScript Hacker News example

LangChain’s community integration includes HNLoader for a Hacker News item. The integration uses Cheerio, so install that source dependency as well:

npm install @langchain/community cheerio

import { HNLoader } from "@langchain/community/document_loaders/web/hn";

const loader = new HNLoader("https://news.ycombinator.com/item?id=8863");
const documents = await loader.load();

for (const document of documents) {
  console.log(document.metadata);
  console.log(document.pageContent.slice(0, 500));
}

This is an example of a supported source integration, not a guarantee that arbitrary sites can be extracted with HNLoader or any other single loader.

Build the RAG index separately from answering

1. Load source data

Run loaders in an indexing job, not inside every user request. Save the source URL, title, fetch time, and any source identifier in document metadata so you can trace an answer back to its input.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Split large documents

Embedding an entire long page produces poor retrieval granularity. Split it into overlapping chunks so a relevant section can be selected without losing nearby context:

from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=150,
)
chunks = splitter.split_documents(documents)
print(f"Indexed {len(chunks)} chunks")

Chunk size and overlap are application settings, not universal constants. Evaluate them against your pages and questions. Preserve headings or other metadata where your loader supplies them.

3. Embed and store chunks

Each chunk is converted to a vector and stored with its text and metadata. The following uses an OpenAI embedding class and Chroma; substitute providers or stores through LangChain’s corresponding integrations:

from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma

embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vector_store = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    collection_name="web-knowledge",
    persist_directory="./chroma-data",
)
retriever = vector_store.as_retriever(search_kwargs={"k": 4})

Keep the indexing job deterministic enough to rerun. If a page changes, replace or version its previous chunks rather than silently accumulating duplicates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieve context and generate an answer

At question time, search the index and pass only the returned context to the model. This avoids fetching an entire site for every question.

from langchain_openai import ChatOpenAI

question = "How does the deployment process work?"
relevant_docs = retriever.invoke(question)
context = "nn".join(
    f"Source: {doc.metadata.get('source', 'unknown')}n{doc.page_content}"
    for doc in relevant_docs
)

prompt = f"""Answer using only the supplied context. If it is insufficient, say so.

Context:
{context}

Question: {question}
"""
model = ChatOpenAI(model="gpt-4o-mini")
answer = model.invoke(prompt)
print(answer.content)

The model name above is an example; select a model available to your account and keep provider-specific configuration outside source code. A production prompt should require the model to distinguish evidence from inference and preserve source metadata for citations in your user interface.

Turn retrieval into an agent tool

LangChain defines an agent as a model calling tools in a loop until the task is complete. The prompt, tools, and middleware form the agent’s harness. create_agent is the configurable entry point; LangChain’s agent implementations use LangGraph primitives, and direct LangGraph construction is the deeper customization route.

Expose the retriever as a narrow tool rather than giving the agent unrestricted database access:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from langchain.agents import create_agent
from langchain.tools import tool

@tool
def search_web_knowledge(query: str) -> str:
    """Search the indexed web documents and return relevant passages."""
    docs = retriever.invoke(query)
    if not docs:
        return "No matching documents were found."
    return "nn".join(
        f"Source: {d.metadata.get('source', 'unknown')}n{d.page_content}"
        for d in docs
    )

agent = create_agent(
    model=model,
    tools=[search_web_knowledge],
    system_prompt=(
        "Use search_web_knowledge when the answer depends on the indexed sources. "
        "Do not invent evidence; state when the sources are insufficient."
    ),
)

result = agent.invoke({
    "messages": [
        {"role": "user", "content": "Compare the deployment options in the indexed guides."}
    ]
})
print(result["messages"][-1].content)

Agent APIs and package layouts are time-sensitive. Check the installed LangChain version’s signature for create_agent, tool decorators, message formats, and middleware before copying this into a pinned application.

Web-ingestion options that affect quality

Dynamic pages and consent UI

A plain HTTP loader may receive a consent wall, bot challenge, or JavaScript shell instead of article text. Choose an integration that can render or otherwise access the source you need, and record failed or empty loads instead of embedding them.

Freshness

Separate scheduled indexing from user queries. Store fetch timestamps and source URLs, then re-index according to how quickly the source changes. A short-lived cache can reduce repeated downloads; a full rebuild is simpler when the corpus is small.

Access control

Do not place private documents in a shared vector store without tenant filtering. Propagate authorization metadata into every chunk and apply the filter before context reaches the model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieval quality

Inspect retrieved chunks, not just final answers. Increase or decrease k, adjust chunk boundaries, preserve headings, and test questions that require information from different pages. If retrieval returns irrelevant passages, changing the model alone will not solve the indexing problem.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting checklist

The loader returns no useful text

  • Print page_content and metadata immediately after load().
  • Check whether the URL serves an interstitial, consent screen, or client-rendered shell.
  • Use the loader documented for that source and install its source-specific dependency, such as Cheerio for the Hacker News integration.

Import or package errors

  • Confirm that the integration package is installed separately from the core package.
  • Compare your import path with the documentation for the exact installed version; LangChain has moved integrations between packages over time.
  • Regenerate a clean virtual environment or lockfile when transitive dependencies conflict.

Answers ignore the documents

  • Log the retrieved documents and verify that they contain the answer.
  • Reduce irrelevant context, set a deliberate k, and instruct the model to say when evidence is missing.
  • Check that the embedding model used for queries is compatible with the one used for indexing.

The agent loops or takes too long

  • Give each tool a precise description and return compact results.
  • Set execution limits in the agent or its runtime and log every tool call.
  • Use two-step RAG when retrieval is mandatory; an agent adds no value to a fixed sequence.

Results are stale or duplicated

  • Attach a canonical source identifier and fetch timestamp to each document.
  • Delete or replace old chunks when a page is re-indexed.
  • Make indexing idempotent so retries do not create another copy of the same page.

Performance, reliability, and cost considerations

Indexing spends work on downloading pages, splitting text, and generating embeddings; query-time work spends retrieval and model calls. Batch indexing, reuse embeddings, and avoid re-fetching unchanged sources. Retrieval latency also depends on the vector database and network, while agent latency varies with the number of model-tool turns. There is no universal speed advantage: the architecture comparison is a control-and-flexibility trade-off.

For reliability, record loader failures, empty documents, embedding errors, retrieval counts, and final source identifiers. Retry transient network operations with limits, but do not retry permanent access denials indefinitely. Keep secrets such as model and provider keys in environment variables. Test the complete path with representative pages, changed pages, malformed pages, and questions for which the correct response is “not enough information.”

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server when visual capture is the missing ingestion step. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be switched off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page and element captures, device and viewport settings, custom CSS and JavaScript, waiting conditions, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names also match those used by other screenshot APIs, which helps when switching.

Use the ScreenshotNeo documentation for the current options. This cURL call saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://news.ycombinator.com 
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://news.ycombinator.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://news.ycombinator.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

Frequently Asked Questions

How can I verify a page before embedding it?

Inspect the loader’s first Document objects, including page_content and metadata, and reject interstitials, empty shells, or unrelated navigation before the splitter runs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should an agent do when retrieval finds nothing?

Return an explicit insufficiency message and ask for clarification or another source; do not let the model fill the gap with an unsupported answer.

Can I change vector stores later?

Yes. LangChain’s retrieval components are modular, so a different vector-store integration can replace the current one while the loader, splitter, and application flow remain separate.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.