Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use LangChain to turn web pages into searchable Document objects, index those documents for retrieval-augmented generation (RAG), and—only when the workflow needs dynamic decisions—let an agent choose retrieval or other tools. A dependable implementation keeps ingestion, indexing, question-time retrieval, and agent orchestration as separate layers.
What LangChain contributes to a web-knowledge workflow
LangChain treats retrieval as runtime access to information that a language model does not reliably contain, including private, recent, or application-specific data. Its retrieval components are modular: document loaders ingest sources, text splitters create searchable chunks, embedding models convert chunks to vectors, vector stores persist those vectors, and retrievers return relevant chunks for a question.
That modularity lets you replace a loader, splitter, embedding provider, or vector store without redesigning the whole application. It does not make every website scrapeable with one universal connector. A loader is an ingestion interface; the extraction mechanism and dependencies depend on the source integration.
Choose two-step or agentic RAG first
The most important design decision is whether retrieval is always required or whether the model should decide when to retrieve and which tool to use.
#1 Best Overall
| Decision axis | Two-step RAG | Agentic RAG |
|---|---|---|
| Retrieval timing | Always before generation | The agent chooses when and how to retrieve |
| Control | Higher | Lower |
| Flexibility | Lower | Higher |
| Latency profile | Generally more predictable | Variable |
| Good fit | FAQs and documentation bots | Research assistants using multiple tools |
Use two-step RAG when every question should search the same knowledge base. It is easier to test, cache, and operate. Use agentic RAG when a task may require several tools, conditional searches, calculations, or a decision about whether retrieval is needed. The additional flexibility brings a less predictable number of model and tool calls. A hybrid can add validation after an agent gathers evidence.
Ingest web pages as documents
Python loader example
Install the packages required by your selected integrations. Package names and import paths evolve, so pin and verify versions in your project before deployment.
pip install -U langchain langchain-community langchain-openai langchain-chroma beautifulsoup4
This example uses the community web loader, then prints the extracted text and metadata before indexing it:
from langchain_community.document_loaders import WebBaseLoader
urls = [
"https://example.com/guide",
"https://example.com/reference",
]
loader = WebBaseLoader(urls)
documents = loader.load()
for document in documents:
print(document.metadata)
print(document.page_content[:500])
Inspecting the result matters. If the page is mostly navigation, an interstitial, or an empty shell rendered by JavaScript, fix extraction before embedding. The loader’s behavior is source-specific; do not assume the same code extracts every site.
Official JavaScript Hacker News example
LangChain’s community integration includes HNLoader for a Hacker News item. The integration uses Cheerio, so install that source dependency as well:
Rank #2
npm install @langchain/community cheerio
import { HNLoader } from "@langchain/community/document_loaders/web/hn";
const loader = new HNLoader("https://news.ycombinator.com/item?id=8863");
const documents = await loader.load();
for (const document of documents) {
console.log(document.metadata);
console.log(document.pageContent.slice(0, 500));
}
This is an example of a supported source integration, not a guarantee that arbitrary sites can be extracted with HNLoader or any other single loader.
Build the RAG index separately from answering
1. Load source data
Run loaders in an indexing job, not inside every user request. Save the source URL, title, fetch time, and any source identifier in document metadata so you can trace an answer back to its input.
2. Split large documents
Embedding an entire long page produces poor retrieval granularity. Split it into overlapping chunks so a relevant section can be selected without losing nearby context:
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=150,
)
chunks = splitter.split_documents(documents)
print(f"Indexed {len(chunks)} chunks")
Chunk size and overlap are application settings, not universal constants. Evaluate them against your pages and questions. Preserve headings or other metadata where your loader supplies them.
3. Embed and store chunks
Each chunk is converted to a vector and stored with its text and metadata. The following uses an OpenAI embedding class and Chroma; substitute providers or stores through LangChain’s corresponding integrations:
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vector_store = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
collection_name="web-knowledge",
persist_directory="./chroma-data",
)
retriever = vector_store.as_retriever(search_kwargs={"k": 4})
Keep the indexing job deterministic enough to rerun. If a page changes, replace or version its previous chunks rather than silently accumulating duplicates.
Recommended Free Tools
Retrieve context and generate an answer
At question time, search the index and pass only the returned context to the model. This avoids fetching an entire site for every question.
from langchain_openai import ChatOpenAI
question = "How does the deployment process work?"
relevant_docs = retriever.invoke(question)
context = "nn".join(
f"Source: {doc.metadata.get('source', 'unknown')}n{doc.page_content}"
for doc in relevant_docs
)
prompt = f"""Answer using only the supplied context. If it is insufficient, say so.
Context:
{context}
Question: {question}
"""
model = ChatOpenAI(model="gpt-4o-mini")
answer = model.invoke(prompt)
print(answer.content)
The model name above is an example; select a model available to your account and keep provider-specific configuration outside source code. A production prompt should require the model to distinguish evidence from inference and preserve source metadata for citations in your user interface.
Turn retrieval into an agent tool
LangChain defines an agent as a model calling tools in a loop until the task is complete. The prompt, tools, and middleware form the agent’s harness. create_agent is the configurable entry point; LangChain’s agent implementations use LangGraph primitives, and direct LangGraph construction is the deeper customization route.
Expose the retriever as a narrow tool rather than giving the agent unrestricted database access:
from langchain.agents import create_agent
from langchain.tools import tool
@tool
def search_web_knowledge(query: str) -> str:
"""Search the indexed web documents and return relevant passages."""
docs = retriever.invoke(query)
if not docs:
return "No matching documents were found."
return "nn".join(
f"Source: {d.metadata.get('source', 'unknown')}n{d.page_content}"
for d in docs
)
agent = create_agent(
model=model,
tools=[search_web_knowledge],
system_prompt=(
"Use search_web_knowledge when the answer depends on the indexed sources. "
"Do not invent evidence; state when the sources are insufficient."
),
)
result = agent.invoke({
"messages": [
{"role": "user", "content": "Compare the deployment options in the indexed guides."}
]
})
print(result["messages"][-1].content)
Agent APIs and package layouts are time-sensitive. Check the installed LangChain version’s signature for create_agent, tool decorators, message formats, and middleware before copying this into a pinned application.
Web-ingestion options that affect quality
Dynamic pages and consent UI
A plain HTTP loader may receive a consent wall, bot challenge, or JavaScript shell instead of article text. Choose an integration that can render or otherwise access the source you need, and record failed or empty loads instead of embedding them.
Freshness
Separate scheduled indexing from user queries. Store fetch timestamps and source URLs, then re-index according to how quickly the source changes. A short-lived cache can reduce repeated downloads; a full rebuild is simpler when the corpus is small.
Access control
Do not place private documents in a shared vector store without tenant filtering. Propagate authorization metadata into every chunk and apply the filter before context reaches the model.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsRetrieval quality
Inspect retrieved chunks, not just final answers. Increase or decrease k, adjust chunk boundaries, preserve headings, and test questions that require information from different pages. If retrieval returns irrelevant passages, changing the model alone will not solve the indexing problem.
Best Value
Troubleshooting checklist
The loader returns no useful text
- Print
page_contentand metadata immediately afterload(). - Check whether the URL serves an interstitial, consent screen, or client-rendered shell.
- Use the loader documented for that source and install its source-specific dependency, such as Cheerio for the Hacker News integration.
Import or package errors
- Confirm that the integration package is installed separately from the core package.
- Compare your import path with the documentation for the exact installed version; LangChain has moved integrations between packages over time.
- Regenerate a clean virtual environment or lockfile when transitive dependencies conflict.
Answers ignore the documents
- Log the retrieved documents and verify that they contain the answer.
- Reduce irrelevant context, set a deliberate
k, and instruct the model to say when evidence is missing. - Check that the embedding model used for queries is compatible with the one used for indexing.
The agent loops or takes too long
- Give each tool a precise description and return compact results.
- Set execution limits in the agent or its runtime and log every tool call.
- Use two-step RAG when retrieval is mandatory; an agent adds no value to a fixed sequence.
Results are stale or duplicated
- Attach a canonical source identifier and fetch timestamp to each document.
- Delete or replace old chunks when a page is re-indexed.
- Make indexing idempotent so retries do not create another copy of the same page.
Performance, reliability, and cost considerations
Indexing spends work on downloading pages, splitting text, and generating embeddings; query-time work spends retrieval and model calls. Batch indexing, reuse embeddings, and avoid re-fetching unchanged sources. Retrieval latency also depends on the vector database and network, while agent latency varies with the number of model-tool turns. There is no universal speed advantage: the architecture comparison is a control-and-flexibility trade-off.
For reliability, record loader failures, empty documents, embedding errors, retrieval counts, and final source identifiers. Retry transient network operations with limits, but do not retry permanent access denials indefinitely. Keep secrets such as model and provider keys in environment variables. Test the complete path with representative pages, changed pages, malformed pages, and questions for which the correct response is “not enough information.”
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server when visual capture is the missing ingestion step. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be switched off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
One GET request returns PNG, JPEG, WebP, or PDF. The API supports full-page and element captures, device and viewport settings, custom CSS and JavaScript, waiting conditions, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names also match those used by other screenshot APIs, which helps when switching.
Use the ScreenshotNeo documentation for the current options. This cURL call saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot"
-d access_key=YOUR_API_KEY
--data-urlencode url=https://news.ycombinator.com
-o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://news.ycombinator.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://news.ycombinator.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.
Frequently Asked Questions
How can I verify a page before embedding it?
Inspect the loader’s first Document objects, including page_content and metadata, and reject interstitials, empty shells, or unrelated navigation before the splitter runs.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallWhat should an agent do when retrieval finds nothing?
Return an explicit insufficiency message and ask for clarification or another source; do not let the model fill the gap with an unsupported answer.
Can I change vector stores later?
Yes. LangChain’s retrieval components are modular, so a different vector-store integration can replace the current one while the loader, splitter, and application flow remain separate.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

