A web crawler discovers URLs, fetches selected pages, and may follow links to find more. To crawl a small site yourself, start with a seed URL, keep a queue and a set of visited URLs, stay within a defined scope, and make requests slowly. Crawling is not the same as indexing: a search engine can fetch a page without storing it in its index or showing it in results.
What is a web crawler?
A web crawler—also called a bot, robot, or spider—is software that automatically discovers and retrieves web resources. There is no central registry of every page on the web. Search engines find URLs from pages they already know, links on those pages, and submitted sitemaps. They then decide which discovered URLs to fetch.
Keep three stages distinct: crawling is fetching a resource; indexing is processing and potentially storing information about it; and serving is deciding whether and how to show it in search results. A successful fetch does not guarantee indexing or a search listing.
How a crawler works
- Choose seed URLs. These are the starting pages, supplied directly or discovered through another source.
- Queue eligible URLs. Track URLs waiting to be fetched and separately track those already seen.
- Fetch a URL. Request its resource, inspect the response, and limit request load.
- Parse the response. Extract the information the task needs, such as links in HTML.
- Normalize and filter links. Resolve relative links, remove fragments where appropriate, deduplicate, enforce scope and access rules, then queue eligible URLs.
- Stop deliberately. Finish when the queue is empty or a page limit, time limit, or crawl boundary is reached.
This is a useful implementation model, not a universal specification. Production crawlers differ in how they schedule requests, parse and render pages, and store results.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
How to crawl a small website with Python
This standard-library example crawls same-host HTML pages from one starting URL. It follows robots.txt rules using Python’s urllib.robotparser, makes one request at a time, waits between requests, and stops at a page cap. It is a teaching example, not a production crawler: it does not render JavaScript, retry transient errors, or save page content.
- Save as
crawler.py. SetSTART_URLto a site you are permitted to crawl. The example uses a reserved example domain; replace it with your target. - Run with Python 3:
python crawler.py.
from collections import deque
from html.parser import HTMLParser
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit
from urllib.robotparser import RobotFileParser
from urllib.request import Request, urlopen
import time
START_URL = "https://example.com/"
USER_AGENT = "LearningCrawler/1.0 (contact: [email protected])"
MAX_PAGES = 50
DELAY_SECONDS = 1.0
TIMEOUT_SECONDS = 15
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag.lower() == "a":
href = dict(attrs).get("href")
if href:
self.links.append(href)
def origin(url):
parts = urlsplit(url)
return f"{parts.scheme}://{parts.netloc}"
start = urldefrag(START_URL)[0]
site_origin = origin(start)
robots_url = site_origin + "/robots.txt"
robot_parser = RobotFileParser(robots_url)
try:
robot_parser.read()
except (OSError, URLError):
print(f"Could not read {robots_url}; stopping rather than assuming permission.")
raise SystemExit(1)
queue = deque([start])
seen = {start}
fetched = 0
while queue and fetched < MAX_PAGES:
url = queue.popleft()
if not robot_parser.can_fetch(USER_AGENT, url):
print(f"SKIP robots.txt: {url}")
continue
request = Request(url, headers={"User-Agent": USER_AGENT})
try:
with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
content_type = response.headers.get("Content-Type", "")
status = response.status
final_url = urldefrag(response.geturl())[0]
body = response.read()
except HTTPError as error:
print(f"HTTP {error.code}: {url}")
time.sleep(DELAY_SECONDS)
continue
except (URLError, TimeoutError, OSError) as error:
print(f"FETCH ERROR: {url}: {error}")
time.sleep(DELAY_SECONDS)
continue
fetched += 1
print(f"{status} {final_url} ({content_type})")
time.sleep(DELAY_SECONDS)
if "text/html" not in content_type.lower():
continue
parser = LinkParser()
try:
parser.feed(body.decode("utf-8", errors="replace"))
except Exception as error:
print(f"PARSE ERROR: {final_url}: {error}")
continue
for href in parser.links:
candidate = urldefrag(urljoin(final_url, href))[0]
parts = urlsplit(candidate)
if parts.scheme not in ("http", "https"):
continue
if origin(candidate) != site_origin or candidate in seen:
continue
seen.add(candidate)
queue.append(candidate)
print(f"Finished: fetched {fetched} page(s); {len(queue)} URL(s) remain queued.")
What this example does not do
- It limits scope to the exact starting origin, so a different subdomain is out of scope.
- It removes URL fragments, but does not canonicalize every equivalent URL or remove query parameters. Sites can expose many distinct URLs for effectively duplicate content.
- It checks robots.txt before each fetch, but robots.txt is not permission to access private material and is not an access-control mechanism.
- It does not honor a universal crawl rate—none exists. The one-second delay is a conservative example setting, not a rule that fits every site.
- It parses links present in the fetched HTML. Content inserted only after JavaScript runs may not be found.
How to crawl responsibly
Identify your crawler with a meaningful user-agent, keep concurrency low, and use delays or backoff when a site responds slowly or with server errors. Search crawlers also manage their request pace; server responses such as HTTP 500 errors can lead Google’s crawler to slow down. Do not treat a successful request as a reason to increase load indefinitely.
What robots.txt controls
A site’s /robots.txt file expresses which paths compliant crawlers may access. For Google, the file is placed at the top level and its rules apply only to the matching host, protocol, and port. Supported directives include user-agent, allow, disallow, and sitemap; Google does not support crawl-delay. Other crawlers may interpret rules differently, so consult the crawler’s documentation.
Robots rules are not security. A blocked URL may still appear in search results if other pages link to it, even if the crawler cannot fetch its content. Protect private pages with authentication or another access-control mechanism. If eligible content should not appear in Google Search, use a mechanism such as noindex or password protection rather than relying on robots.txt alone.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #3
Links and sitemaps
Links help crawlers discover pages as they traverse known pages. An XML sitemap can also list URLs for a crawler to consider, but inclusion is not a guarantee that a URL will be fetched or indexed. If you use a sitemap, keep it current; for updated content, include an accurate lastmod value.
What crawl budget means
Google describes crawl budget as the set of URLs it can and wants to crawl. Capacity concerns how much fetching a host can tolerate; demand reflects which URLs Google considers worth revisiting or discovering. Demand can vary with site size, update frequency, page quality, relevance, popularity, URL inventory, and how stale content may be. There is no single universal crawl rate or threshold for every site.
Reduce wasted crawling
- Consolidate duplicate pages and avoid generating unnecessary URL variants.
- Limit unbounded combinations from filters, sorting, faceted navigation, calendars, and session IDs.
- Keep sitemaps current and avoid long redirect chains.
- Return 404 or 410 for pages that have been permanently removed.
- Check relative links for malformed paths that can create unintended URL spaces.
When JavaScript rendering matters
A simple crawler can retrieve HTML and extract links without launching a browser. That is often enough for basic link discovery, but it may miss content or links added only after client-side JavaScript executes. Google’s crawler renders pages and runs JavaScript; a custom crawler may or may not need to. Use browser rendering when the pages or task require it, because it adds processing cost and implementation complexity.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your task is to capture a clean screenshot of a page rather than build a link-following crawler, ScreenshotNeo provides a screenshot API and MCP server. This is not a substitute for a crawler that discovers and traverses URLs; it is a one-request way to capture a known URL as an image or PDF.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
For API options and setup, see the ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
- Before capture, it accepts cookie/consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off.
- Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Frequently Asked Questions
Does a page fetched by a crawler automatically appear in Google Search?
No. Fetching, indexing, and serving search results are separate stages; crawling alone does not guarantee inclusion.
Can a basic Python crawler see every page on a JavaScript-heavy site?
No. It only parses links in the HTML it fetches. Pages or links created after JavaScript executes may require a rendering-capable crawler.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




