Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk6 min

How to Crawl Websites with Python: A Practical Scrapy Guide

Learn when urllib is enough and how to use Scrapy to follow links, extract structured data, export results, and keep a Python crawl in scope.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For one page, Python’s urllib.request.urlopen() can fetch the response. To crawl multiple pages—follow links, extract structured data, and export results—Scrapy provides the scheduling and spider workflow you would otherwise have to build yourself. This guide starts with a one-page fetch, then builds a small Scrapy crawler and explains how to keep its scope and request behavior appropriate for the site.

Fetch one page with Python

A single request is not yet a crawler, but it is a useful way to retrieve a page when you already know its URL. Python’s urllib.request HOWTO documents urlopen() for opening a URL and reading its response. See the Python 3.14.7 urllib HOWTO.

from urllib.request import urlopen

url = "https://example.com/"
with urlopen(url, timeout=20) as response:
    html = response.read()

print(html[:500])

The result is response bytes, not a structured extraction of a page. This small interface does not supply a crawl scheduler, link-following rules, or a feed-export workflow; you would need to implement those pieces yourself if the task grows.

Use Scrapy for a multi-page crawl

Scrapy spiders generate requests and process responses. A callback can yield extracted items and schedule additional requests, making it a better fit when you want to traverse a site and save structured results. Its framework includes a scheduler, downloader, spiders, items, pipelines, and feed exports. See the Scrapy overview.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

1. Create a project and identify your crawler

Install Scrapy in your Python environment, then create a project from the command line:

python -m pip install Scrapy
scrapy startproject sitecrawl
cd sitecrawl

Set a descriptive, project-specific user agent in sitecrawl/settings.py. Include an operator contact URL or email that you control; do not copy a fictitious contact address. Site owners can use an identifiable user agent to ask you to adjust the crawler if needed. The official Scrapy tutorial demonstrates this practice.

USER_AGENT = "SiteCrawl (+https://your-domain.example/contact)"

Replace the example contact URL with a real one before crawling. The setting identifies your crawler; it does not grant permission to access a site.

2. Write a spider for the target’s actual structure

Create sitecrawl/spiders/articles.py. This example starts from a known listing page, extracts article titles and links, and schedules requests for linked article pages on the same host. Selectors are examples: inspect the target HTML and change them to match its markup.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import scrapy
from urllib.parse import urlparse


class ArticlesSpider(scrapy.Spider):
    name = "articles"
    start_urls = ["https://example.com/articles/"]
    allowed_domains = ["example.com"]

    def parse(self, response):
        for link in response.css("a.article-link"):
            href = link.attrib.get("href")
            title = " ".join(link.css("::text").getall()).strip()
            if href:
                yield response.follow(
                    href,
                    callback=self.parse_article,
                    meta={"listing_title": title},
                )

        next_page = response.css("a.next::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

    def parse_article(self, response):
        yield {
            "url": response.url,
            "title": response.css("h1::text").get(default="").strip(),
            "listing_title": response.meta.get("listing_title", ""),
            "text": " ".join(
                part.strip()
                for part in response.css("article p::text").getall()
                if part.strip()
            ),
        }

allowed_domains helps constrain the crawl to the intended host, while response.follow() resolves relative links against the page being parsed. The selectors a.article-link, a.next, and article p must fit the site; they are not universal page selectors. Scrapy’s tutorial shows the callback-and-yield pattern for extracting data and following links.

3. Run the spider and export data

From the project directory, run the spider and export its items as JSON Lines:

scrapy crawl articles -O articles.jsonl

Each yielded item is written as a separate JSON object. The tutorial demonstrates command-line feed export; for a larger workflow, Scrapy item pipelines can validate, clean, or store items, and feeds can be sent to different destinations.

Choose the right link-discovery pattern

Use the simplest spider type that matches the site’s structure and your extraction needs. Scrapy documents plain spiders, rule-based crawling, and sitemap-based discovery; its spider documentation notes that CrawlSpider is not suitable for every site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach Best fit Trade-off
Plain Spider Custom traversal, unusual page structure, or bespoke extraction You control the callbacks and logic, and must maintain them.
CrawlSpider A regular site whose links can be expressed as rules Convenient rule-based following, but rules and custom callbacks require care and may not fit every site.
SitemapSpider A site with useful sitemap URLs Discovers URLs from sitemap structure rather than relying only on links found on pages.

Decide based on the site’s link structure, whether a usable sitemap exists, the data you need, and how you will constrain requests. Do not choose a crawler on the assumption that maximum request speed is the goal.

Set scope and crawl responsibly

Check robots.txt and site requirements

Before crawling, inspect the target’s top-level /robots.txt file and configure your crawler to respect the applicable instructions. RFC 9309 defines the Robots Exclusion Protocol and specifies the top-level path for the file; Scrapy lists robots.txt support in its settings documentation. A robots file is not a substitute for reviewing the site’s terms or applicable law. The protocol describes crawler instructions; it does not determine legal permission for a particular site, data type, or jurisdiction. See RFC 9309.

Limit what the spider can reach

  • Start from the smallest set of URLs that answers your question.
  • Use an appropriate domain scope and link rules so the spider does not wander into unrelated sections or external sites.
  • Review pagination, query parameters, calendars, and other URL patterns that can create very large or repetitive crawl spaces.
  • Use an identifiable user agent and a request pace appropriate to the target. Scrapy supports concurrent requests and crawl-politeness controls; configure them for the site instead of treating maximum speed as the goal.

Site structure and response behavior vary, so no generic spider can guarantee it will reach every page. A crawl should be treated as a defined collection process, not proof that every page was discovered.

Troubleshoot common crawl problems

  • No items are exported: Confirm the spider name and command, then inspect the response and update CSS selectors to match the actual markup. The example selectors are placeholders for site-specific structure.
  • Only listing pages are fetched: Check that the listing selector matches article links, that links have an href, and that the callback yielding article requests is reached.
  • The spider leaves the intended scope: Tighten allowed_domains and review follow rules, pagination, and parameterized URLs.
  • Requests are denied or behavior changes: Review the site’s robots instructions and terms, reduce request intensity where appropriate, and use a descriptive user agent with a real operator contact.
  • The crawl produces duplicates or excessive URLs: Inspect link patterns such as tracking or filter parameters and refine the rules or URL handling to match the target’s structure.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a screenshot is the output you need

A crawler extracts page data; it is not the right output when you need a visual record of a rendered page. For a screenshot or PDF through one API request, ScreenshotNeo offers a website screenshot API and MCP server for developers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

For a screenshot, make a GET request with the target URL. The API returns PNG, JPEG, WebP, or PDF output. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response indicates the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card.

Version context

The Scrapy project page reports version 2.19.0, released in September 2026. Commands, defaults, and settings can change, so consult the current Scrapy documentation for the version you install. The Python fetch example follows the Python 3.14.7 urllib HOWTO.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.