For one page, Python’s urllib.request.urlopen() can fetch the response. To crawl multiple pages—follow links, extract structured data, and export results—Scrapy provides the scheduling and spider workflow you would otherwise have to build yourself. This guide starts with a one-page fetch, then builds a small Scrapy crawler and explains how to keep its scope and request behavior appropriate for the site.
Fetch one page with Python
A single request is not yet a crawler, but it is a useful way to retrieve a page when you already know its URL. Python’s urllib.request HOWTO documents urlopen() for opening a URL and reading its response. See the Python 3.14.7 urllib HOWTO.
from urllib.request import urlopen
url = "https://example.com/"
with urlopen(url, timeout=20) as response:
html = response.read()
print(html[:500])
The result is response bytes, not a structured extraction of a page. This small interface does not supply a crawl scheduler, link-following rules, or a feed-export workflow; you would need to implement those pieces yourself if the task grows.
Use Scrapy for a multi-page crawl
Scrapy spiders generate requests and process responses. A callback can yield extracted items and schedule additional requests, making it a better fit when you want to traverse a site and save structured results. Its framework includes a scheduler, downloader, spiders, items, pipelines, and feed exports. See the Scrapy overview.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
1. Create a project and identify your crawler
Install Scrapy in your Python environment, then create a project from the command line:
python -m pip install Scrapy
scrapy startproject sitecrawl
cd sitecrawl
Set a descriptive, project-specific user agent in sitecrawl/settings.py. Include an operator contact URL or email that you control; do not copy a fictitious contact address. Site owners can use an identifiable user agent to ask you to adjust the crawler if needed. The official Scrapy tutorial demonstrates this practice.
USER_AGENT = "SiteCrawl (+https://your-domain.example/contact)"
Replace the example contact URL with a real one before crawling. The setting identifies your crawler; it does not grant permission to access a site.
Rank #2
2. Write a spider for the target’s actual structure
Create sitecrawl/spiders/articles.py. This example starts from a known listing page, extracts article titles and links, and schedules requests for linked article pages on the same host. Selectors are examples: inspect the target HTML and change them to match its markup.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
import scrapy
from urllib.parse import urlparse
class ArticlesSpider(scrapy.Spider):
name = "articles"
start_urls = ["https://example.com/articles/"]
allowed_domains = ["example.com"]
def parse(self, response):
for link in response.css("a.article-link"):
href = link.attrib.get("href")
title = " ".join(link.css("::text").getall()).strip()
if href:
yield response.follow(
href,
callback=self.parse_article,
meta={"listing_title": title},
)
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
def parse_article(self, response):
yield {
"url": response.url,
"title": response.css("h1::text").get(default="").strip(),
"listing_title": response.meta.get("listing_title", ""),
"text": " ".join(
part.strip()
for part in response.css("article p::text").getall()
if part.strip()
),
}
allowed_domains helps constrain the crawl to the intended host, while response.follow() resolves relative links against the page being parsed. The selectors a.article-link, a.next, and article p must fit the site; they are not universal page selectors. Scrapy’s tutorial shows the callback-and-yield pattern for extracting data and following links.
3. Run the spider and export data
From the project directory, run the spider and export its items as JSON Lines:
scrapy crawl articles -O articles.jsonl
Each yielded item is written as a separate JSON object. The tutorial demonstrates command-line feed export; for a larger workflow, Scrapy item pipelines can validate, clean, or store items, and feeds can be sent to different destinations.
Choose the right link-discovery pattern
Use the simplest spider type that matches the site’s structure and your extraction needs. Scrapy documents plain spiders, rule-based crawling, and sitemap-based discovery; its spider documentation notes that CrawlSpider is not suitable for every site.
| Approach | Best fit | Trade-off |
|---|---|---|
Plain Spider |
Custom traversal, unusual page structure, or bespoke extraction | You control the callbacks and logic, and must maintain them. |
CrawlSpider |
A regular site whose links can be expressed as rules | Convenient rule-based following, but rules and custom callbacks require care and may not fit every site. |
SitemapSpider |
A site with useful sitemap URLs | Discovers URLs from sitemap structure rather than relying only on links found on pages. |
Decide based on the site’s link structure, whether a usable sitemap exists, the data you need, and how you will constrain requests. Do not choose a crawler on the assumption that maximum request speed is the goal.
Set scope and crawl responsibly
Check robots.txt and site requirements
Before crawling, inspect the target’s top-level /robots.txt file and configure your crawler to respect the applicable instructions. RFC 9309 defines the Robots Exclusion Protocol and specifies the top-level path for the file; Scrapy lists robots.txt support in its settings documentation. A robots file is not a substitute for reviewing the site’s terms or applicable law. The protocol describes crawler instructions; it does not determine legal permission for a particular site, data type, or jurisdiction. See RFC 9309.
Limit what the spider can reach
- Start from the smallest set of URLs that answers your question.
- Use an appropriate domain scope and link rules so the spider does not wander into unrelated sections or external sites.
- Review pagination, query parameters, calendars, and other URL patterns that can create very large or repetitive crawl spaces.
- Use an identifiable user agent and a request pace appropriate to the target. Scrapy supports concurrent requests and crawl-politeness controls; configure them for the site instead of treating maximum speed as the goal.
Site structure and response behavior vary, so no generic spider can guarantee it will reach every page. A crawl should be treated as a defined collection process, not proof that every page was discovered.
Troubleshoot common crawl problems
- No items are exported: Confirm the spider name and command, then inspect the response and update CSS selectors to match the actual markup. The example selectors are placeholders for site-specific structure.
- Only listing pages are fetched: Check that the listing selector matches article links, that links have an
href, and that the callback yielding article requests is reached. - The spider leaves the intended scope: Tighten
allowed_domainsand review follow rules, pagination, and parameterized URLs. - Requests are denied or behavior changes: Review the site’s robots instructions and terms, reduce request intensity where appropriate, and use a descriptive user agent with a real operator contact.
- The crawl produces duplicates or excessive URLs: Inspect link patterns such as tracking or filter parameters and refine the rules or URL handling to match the target’s structure.
When a screenshot is the output you need
A crawler extracts page data; it is not the right output when you need a visual record of a rendered page. For a screenshot or PDF through one API request, ScreenshotNeo offers a website screenshot API and MCP server for developers.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesBest Value
Or skip the browser setup
For a screenshot, make a GET request with the target URL. The API returns PNG, JPEG, WebP, or PDF output. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response indicates the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card.
Version context
The Scrapy project page reports version 2.19.0, released in September 2026. Commands, defaults, and settings can change, so consult the current Scrapy documentation for the version you install. The Python fetch example follows the Python 3.14.7 urllib HOWTO.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




