October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk7 min

Web Crawlers Explained: How to Crawl a Website

A web crawler finds URLs, fetches selected pages, and follows links. Learn the crawl loop, build a small Python crawler, and understand robots.txt, sitemaps, crawl budget, and JavaScript rendering.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A web crawler discovers URLs, fetches selected pages, and may follow links to find more. To crawl a small site yourself, start with a seed URL, keep a queue and a set of visited URLs, stay within a defined scope, and make requests slowly. Crawling is not the same as indexing: a search engine can fetch a page without storing it in its index or showing it in results.

What is a web crawler?

A web crawler—also called a bot, robot, or spider—is software that automatically discovers and retrieves web resources. There is no central registry of every page on the web. Search engines find URLs from pages they already know, links on those pages, and submitted sitemaps. They then decide which discovered URLs to fetch.

Keep three stages distinct: crawling is fetching a resource; indexing is processing and potentially storing information about it; and serving is deciding whether and how to show it in search results. A successful fetch does not guarantee indexing or a search listing.

How a crawler works

  1. Choose seed URLs. These are the starting pages, supplied directly or discovered through another source.
  2. Queue eligible URLs. Track URLs waiting to be fetched and separately track those already seen.
  3. Fetch a URL. Request its resource, inspect the response, and limit request load.
  4. Parse the response. Extract the information the task needs, such as links in HTML.
  5. Normalize and filter links. Resolve relative links, remove fragments where appropriate, deduplicate, enforce scope and access rules, then queue eligible URLs.
  6. Stop deliberately. Finish when the queue is empty or a page limit, time limit, or crawl boundary is reached.

This is a useful implementation model, not a universal specification. Production crawlers differ in how they schedule requests, parse and render pages, and store results.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to crawl a small website with Python

This standard-library example crawls same-host HTML pages from one starting URL. It follows robots.txt rules using Python’s urllib.robotparser, makes one request at a time, waits between requests, and stops at a page cap. It is a teaching example, not a production crawler: it does not render JavaScript, retry transient errors, or save page content.

  1. Save as crawler.py. Set START_URL to a site you are permitted to crawl. The example uses a reserved example domain; replace it with your target.
  2. Run with Python 3: python crawler.py.
from collections import deque
from html.parser import HTMLParser
from urllib.error import HTTPError, URLError
from urllib.parse import urldefrag, urljoin, urlsplit
from urllib.robotparser import RobotFileParser
from urllib.request import Request, urlopen
import time

START_URL = "https://example.com/"
USER_AGENT = "LearningCrawler/1.0 (contact: [email protected])"
MAX_PAGES = 50
DELAY_SECONDS = 1.0
TIMEOUT_SECONDS = 15

class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag.lower() == "a":
            href = dict(attrs).get("href")
            if href:
                self.links.append(href)

def origin(url):
    parts = urlsplit(url)
    return f"{parts.scheme}://{parts.netloc}"

start = urldefrag(START_URL)[0]
site_origin = origin(start)
robots_url = site_origin + "/robots.txt"
robot_parser = RobotFileParser(robots_url)
try:
    robot_parser.read()
except (OSError, URLError):
    print(f"Could not read {robots_url}; stopping rather than assuming permission.")
    raise SystemExit(1)

queue = deque([start])
seen = {start}
fetched = 0

while queue and fetched < MAX_PAGES:
    url = queue.popleft()
    if not robot_parser.can_fetch(USER_AGENT, url):
        print(f"SKIP robots.txt: {url}")
        continue

    request = Request(url, headers={"User-Agent": USER_AGENT})
    try:
        with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            content_type = response.headers.get("Content-Type", "")
            status = response.status
            final_url = urldefrag(response.geturl())[0]
            body = response.read()
    except HTTPError as error:
        print(f"HTTP {error.code}: {url}")
        time.sleep(DELAY_SECONDS)
        continue
    except (URLError, TimeoutError, OSError) as error:
        print(f"FETCH ERROR: {url}: {error}")
        time.sleep(DELAY_SECONDS)
        continue

    fetched += 1
    print(f"{status} {final_url} ({content_type})")
    time.sleep(DELAY_SECONDS)

    if "text/html" not in content_type.lower():
        continue

    parser = LinkParser()
    try:
        parser.feed(body.decode("utf-8", errors="replace"))
    except Exception as error:
        print(f"PARSE ERROR: {final_url}: {error}")
        continue

    for href in parser.links:
        candidate = urldefrag(urljoin(final_url, href))[0]
        parts = urlsplit(candidate)
        if parts.scheme not in ("http", "https"):
            continue
        if origin(candidate) != site_origin or candidate in seen:
            continue
        seen.add(candidate)
        queue.append(candidate)

print(f"Finished: fetched {fetched} page(s); {len(queue)} URL(s) remain queued.")

What this example does not do

  • It limits scope to the exact starting origin, so a different subdomain is out of scope.
  • It removes URL fragments, but does not canonicalize every equivalent URL or remove query parameters. Sites can expose many distinct URLs for effectively duplicate content.
  • It checks robots.txt before each fetch, but robots.txt is not permission to access private material and is not an access-control mechanism.
  • It does not honor a universal crawl rate—none exists. The one-second delay is a conservative example setting, not a rule that fits every site.
  • It parses links present in the fetched HTML. Content inserted only after JavaScript runs may not be found.

How to crawl responsibly

Identify your crawler with a meaningful user-agent, keep concurrency low, and use delays or backoff when a site responds slowly or with server errors. Search crawlers also manage their request pace; server responses such as HTTP 500 errors can lead Google’s crawler to slow down. Do not treat a successful request as a reason to increase load indefinitely.

What robots.txt controls

A site’s /robots.txt file expresses which paths compliant crawlers may access. For Google, the file is placed at the top level and its rules apply only to the matching host, protocol, and port. Supported directives include user-agent, allow, disallow, and sitemap; Google does not support crawl-delay. Other crawlers may interpret rules differently, so consult the crawler’s documentation.

Robots rules are not security. A blocked URL may still appear in search results if other pages link to it, even if the crawler cannot fetch its content. Protect private pages with authentication or another access-control mechanism. If eligible content should not appear in Google Search, use a mechanism such as noindex or password protection rather than relying on robots.txt alone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Links and sitemaps

Links help crawlers discover pages as they traverse known pages. An XML sitemap can also list URLs for a crawler to consider, but inclusion is not a guarantee that a URL will be fetched or indexed. If you use a sitemap, keep it current; for updated content, include an accurate lastmod value.

What crawl budget means

Google describes crawl budget as the set of URLs it can and wants to crawl. Capacity concerns how much fetching a host can tolerate; demand reflects which URLs Google considers worth revisiting or discovering. Demand can vary with site size, update frequency, page quality, relevance, popularity, URL inventory, and how stale content may be. There is no single universal crawl rate or threshold for every site.

Reduce wasted crawling

  • Consolidate duplicate pages and avoid generating unnecessary URL variants.
  • Limit unbounded combinations from filters, sorting, faceted navigation, calendars, and session IDs.
  • Keep sitemaps current and avoid long redirect chains.
  • Return 404 or 410 for pages that have been permanently removed.
  • Check relative links for malformed paths that can create unintended URL spaces.

When JavaScript rendering matters

A simple crawler can retrieve HTML and extract links without launching a browser. That is often enough for basic link discovery, but it may miss content or links added only after client-side JavaScript executes. Google’s crawler renders pages and runs JavaScript; a custom crawler may or may not need to. Use browser rendering when the pages or task require it, because it adds processing cost and implementation complexity.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your task is to capture a clean screenshot of a page rather than build a link-following crawler, ScreenshotNeo provides a screenshot API and MCP server. This is not a substitute for a crawler that discovers and traverses URLs; it is a one-request way to capture a known URL as an image or PDF.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For API options and setup, see the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
  • Before capture, it accepts cookie/consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off.
  • Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Frequently Asked Questions

Does a page fetched by a crawler automatically appear in Google Search?

No. Fetching, indexing, and serving search results are separate stages; crawling alone does not guarantee inclusion.

Can a basic Python crawler see every page on a JavaScript-heavy site?

No. It only parses links in the HTML it fetches. Pages or links created after JavaScript executes may require a rendering-capable crawler.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.