Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Build an AI-ready crawler as a permission-aware Scrapy project: decide what you may fetch, identify the crawler clearly, respect robots.txt and rate limits, extract and validate structured records, and preserve provenance before sending anything to search, embeddings, or an LLM. Use a browser only when the required content is missing from the permitted raw response. That approach keeps the crawl easier to debug and helps prevent empty, duplicated, or poorly parsed pages from contaminating downstream systems.
What makes a crawler AI-ready?
A crawler is not AI-ready just because it downloads HTML. It needs to produce clean, useful content with enough metadata to identify its source, determine when it was retrieved, deduplicate it, and refresh or rebuild an index later. Treat every crawled page as a document with provenance—not as an anonymous text blob.
Before coding, define two contracts:
- Crawl contract: approved domains and paths, exclusions, maximum depth, concurrency, delays, retry behavior, language handling, canonicalization, and retention.
- Output contract: the fields every accepted record must contain, how missing or malformed data is represented, and which records are rejected or quarantined.
A practical record can include the requested URL, canonical URL, retrieval timestamp, publication and update dates when available, title, author, site name, language, cleaned Markdown or text, links, HTTP status, content type, parser version, content hash, extraction status, and warnings. Keep headings, lists, tables, code, captions, and link targets when they carry meaning for retrieval or citation.
Do not chunk pages before cleaning and normalizing them. When you do chunk them, carry document-level provenance into every chunk so a search result can be traced to its page and crawl run.
#1 Best Overall
Set access rules before fetching pages
Make permission checks part of the first request path, not a cleanup task after crawling. Fetch and evaluate the target site’s robots.txt before scheduling pages. Use an identifiable user agent, obey applicable disallow rules and crawl delays, and follow the site’s published terms. OpenAI’s guidance describes robots.txt as a way for site owners to tell crawlers whether they may access particular paths: OpenAI Help Center guidance.
Scrapy has robots.txt support and exposes ROBOTSTXT_USER_AGENT; its documented default Protego parser supports wildcard matching and rule precedence. See Scrapy downloader middleware settings. Robots rules are not a license to ignore other access controls: authentication requirements, terms, rate limits, and site-specific policies still matter.
Handle 401, 403, 429, and challenge pages as access or throttling signals. Do not brute-force through them, evade CAPTCHAs, or treat a block as a reason to rotate identities. OpenAI notes that WAFs, CDNs, bot mitigation, JavaScript challenges, CAPTCHAs, authentication, and geographic rules can prevent an otherwise legitimate crawler from accessing a page: its crawler access guidance.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsOpenAI documents OAI-SearchBot and GPTBot as separate controls: OAI-SearchBot is associated with visibility in ChatGPT search, while GPTBot is associated with training use. A site can manage them independently. OpenAI also says robots.txt changes can take about 24 hours to adjust for its search systems. These are OpenAI-specific statements, not universal timing guarantees for every crawler: Overview of OpenAI crawlers.
Build the crawl around Scrapy
Scrapy’s spiders define how a site is scraped: they follow links and return structured items or more requests from callbacks. Its core also provides selectors, duplicate filtering, robots.txt support, and feed exports. See the spider documentation and Scrapy overview.
1. Create a project and install dependencies
Use a supported Python environment, then install Scrapy and Trafilatura. Trafilatura is useful for article-like pages; it is not a universal parser for every page type.
python -m venv .venv
source .venv/bin/activate
python -m pip install scrapy trafilatura
scrapy startproject ai_crawler
cd ai_crawler
Rank #2
On Windows PowerShell, activate the environment with .venvScriptsActivate.ps1 instead of the Unix activation command.
2. Configure the crawl boundary
Edit the generated ai_crawler/settings.py. Replace the example identity with one that accurately identifies your crawler and includes a contact method you actually monitor. Set the allowed domain in the spider as well as the robots and pacing settings here.
BOT_NAME = "ai_crawler"
USER_AGENT = "ai_crawler (+https://example.com/crawler-info)"
ROBOTSTXT_OBEY = True
ROBOTSTXT_USER_AGENT = USER_AGENT
CONCURRENT_REQUESTS = 4
DOWNLOAD_DELAY = 1.0
FEED_EXPORT_ENCODING = "utf-8"
These are conservative example settings, not a universal safe rate. Follow the site’s stated limits and adjust concurrency and delay accordingly. Scrapy’s duplicate filter helps avoid scheduling the same request repeatedly, but canonical URL normalization and content-level deduplication are still application decisions.
3. Define a page-family spider and output item
Start with one page family on one approved domain. The example below only follows links within example.com and extracts common article metadata. The CSS selectors are deliberately generic: inspect permitted pages and adapt them to the site’s actual HTML. Do not assume every site uses the same title, date, or article-body markup.
Create ai_crawler/spiders/site.py:
import hashlib
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
import scrapy
import trafilatura
class SiteSpider(scrapy.Spider):
name = "site"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/"]
def parse(self, response):
content_type = response.headers.get(b"Content-Type", b"").decode("latin-1")
if "html" not in content_type.lower():
return
Free tools Windows power users keep installed
One-click scans. No signup required.
downloaded = trafilatura.bare_extraction(
response.text,
url=response.url,
output_format="markdown",
with_metadata=True,
)
if downloaded:
body = downloaded.get("raw_text") or ""
canonical = response.css('link[rel="canonical"]::attr(href)').get()
canonical_url = urljoin(response.url, canonical) if canonical else response.url
record = {
"url": response.url,
"canonical_url": canonical_url,
"retrieved_at": datetime.now(timezone.utc).isoformat(),
"published_at": downloaded.get("date"),
"updated_at": None,
"title": downloaded.get("title"),
"author": downloaded.get("author"),
"site_name": downloaded.get("sitename"),
"language": downloaded.get("language"),
"content_markdown": body,
"content_hash": hashlib.sha256(body.encode("utf-8")).hexdigest(),
"http_status": response.status,
"content_type": content_type,
"parser_version": "site-parser-1",
"extraction_status": "ok" if body.strip() else "empty_body",
"extraction_warnings": [],
}
yield record
for href in response.css("a::attr(href)").getall():
target = response.urljoin(href)
parsed = urlparse(target)
if parsed.scheme in ("http", "https") and parsed.hostname == "example.com":
yield response.follow(target, callback=self.parse)
Run it and export JSON Lines so each item is one independently processable record:
scrapy crawl site -O pages.jl
The example records retrieval time and extraction status, but a production pipeline should also record redirect details, warnings, and the actual parser build or code version. A literal parser label such as site-parser-1 is only useful if you update it when extraction behavior changes.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →4. Make inclusion, canonicalization, and freshness explicit
Do not let “follow every link” become the crawl policy. Add path allowlists and exclusions for search results, calendars, session URLs, tracking parameters, and other routes that do not represent content you want to index. Normalize URLs according to the site’s behavior: fragments usually do not identify separate server documents, while query parameters may be meaningful. Preserve the fetched URL even when you store a separate canonical URL.
For incremental crawls, define what counts as a changed document. A content hash can identify identical extracted bodies, but it does not replace source URL, retrieval time, publication date, or parser version. Keep those fields to support citation, refresh decisions, and reproducibility.
Extract content that works for retrieval and RAG
Raw HTML includes navigation, ads, banners, repeated headers, and scripts that can crowd out relevant passages. Scrapy’s extraction guide describes using Trafilatura to extract clean text or Markdown and metadata such as title, author, date, and site name. It also warns that article-focused extraction may return little or nothing for product pages or listings: Scrapy’s extraction guide.
Choose extraction by page family. Article pages, product pages, documentation, forums, and listings often need different selectors or parsers. Preserve structural elements when they change meaning: a table without its headers, code without formatting, or a list flattened into one string may be poor retrieval material even if it is technically text.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallKeep the original HTML or a content hash when you need to audit extraction changes. For each normalized record, distinguish a legitimate empty result from a fetch error or parse failure. Do not silently turn extraction failure into a successful blank document.
Use browser rendering only when the response needs it
Before reaching for browser automation, inspect the HTTP response Scrapy receives. The desired data may already be present in HTML, embedded JavaScript state, or a permitted JSON resource. Scrapy’s dynamic-content guidance specifically recommends checking the HTTP-client response before deciding a browser is required: Dynamic content in Scrapy.
Use scrapy-playwright for pages where meaningful content appears only after JavaScript runs, scrolling, interaction, or client-side requests. Keep that subset narrow: browsers increase compute use, operational complexity, and failure modes. Prefer a direct data endpoint or embedded state object when it is available and you are permitted to use it.
A screenshot is a visual capture, not a substitute for a crawler that discovers links, extracts structured fields, and builds records. If the task is specifically to capture a page image or PDF, ScreenshotNeo is a separate option; its website screenshot API is not a general-purpose web-crawling or RAG ingestion system.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Validate records before indexing or prompting a model
Build fixtures for every important template and test representative page variants before sending content to embeddings, search, or an LLM. At minimum, check required fields, title and date parsing, canonical URLs, body length, link extraction, and boilerplate removal. Compare pages over time so a template change does not quietly turn a useful document into a navigation dump.
Quarantine records that fail validation instead of indexing them. Track sudden changes in status codes, empty-body rates, missing-field rates, duplicate ratios, and content-length distributions. Preserve parser version and crawl timestamp so you can identify affected records and rebuild an index after correcting extraction.
Best Value
Scrapy’s documented AI workflow includes defining a schema, downloading several pages, comparing variants, validating the extraction specification, and generating page objects, spiders, and a runnable test suite: Scrapy’s AI workflow.
Plan for failures, cost, and scale
Access and response failures
- 401 or 403: treat these as an authentication or access denial. Confirm permission and credentials through the site’s documented process; do not attempt to bypass controls.
- 429: slow down or pause according to the site’s policy. Do not increase concurrency in response to throttling.
- Challenge or CAPTCHA page: stop and investigate whether the crawler is permitted and configured appropriately. Do not automate a bypass.
- Timeout or transient server error: use bounded retries with a delay and record the eventual status. Avoid retry loops that amplify load.
Extraction failures
- Empty body: check content type, response HTML, and whether extraction is appropriate for that page family. A product listing may need a page-specific parser rather than article extraction.
- Wrong title, date, or canonical URL: compare the selector or extracted metadata against fixtures for multiple page variants; quarantine invalid records until corrected.
- Sudden duplicate growth: inspect URL normalization, pagination, query parameters, and canonicalization rules before indexing.
- Parser drift: alert on field-null rates, body lengths, status distributions, and duplicate ratios; preserve the prior parser version to identify which crawl introduced the change.
Operational cost and reliability
Keep discovery, fetching, extraction, validation, and indexing separable so a failure in one stage can be retried without repeating every other stage. Network volume grows with the number and size of fetched pages; browser rendering adds browser CPU and additional operational surface. Proxy usage, storage, monitoring, and managed deployment can add further cost. Estimate from your own permitted crawl volume and page mix rather than assuming a universal per-page cost.
Scrapy’s site lists optional extensions including browser rendering through scrapy-playwright, Spidermon for monitoring, Zyte API for proxy rotation and ban avoidance, page objects through scrapy-poet, Scrapy Cloud deployment, and an MCP server for inspecting live crawls: Scrapy project site. They are options around the Scrapy core, not prerequisites. Adopt a hosted service only when volume, rendering needs, reliability, or operational capacity justifies it, and check its current terms and compliance requirements before use.
Or skip the browser setup
For a one-call visual capture—not link discovery or structured crawling—ScreenshotNeo accepts a URL and returns a screenshot or PDF. Its API documentation describes the request options.
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with verdict and billing details in response headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The Free plan includes 1,000 screenshots monthly with no card; paid plans start at $5 for 3,000. All features are available on every plan. Sign up for 1,000 free screenshots a month with no card.
FAQ
Does a robots.txt allow rule guarantee a page is okay to crawl?
No. It addresses crawler access to paths, but does not replace site terms, authentication requirements, rate limits, or other applicable access controls.
Should I send HTML or cleaned Markdown to an LLM?
Use normalized, cleaned content for retrieval in most pipelines, while retaining source HTML or a reproducibility aid when you need to audit parsing. Preserve structure and source metadata that affect meaning and citation.
Can a screenshot API build a knowledge base from a whole site?
No. A screenshot API captures a requested page visually; a crawler still needs to discover permitted pages, extract content and metadata, validate records, and manage refreshes.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

