Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Start a Python crawler with requests and Beautiful Soup when the page’s data is already in its HTTP response. Move to Scrapy when you need to schedule and manage a multi-page crawl; use Playwright only when the page depends on JavaScript execution or browser interaction. A browser is not automatically a better scraper: it is slower and more resource-intensive than an HTTP request, and its UI selectors can be fragile.
This tutorial builds those skills in order, from one respectful static-page fetch to pagination, crawl controls, and browser-rendered pages. Use it only where you are allowed to access and collect the information.
Choose the right Python tool for the page
| Tool | What it does | Best fit |
|---|---|---|
requests |
Fetches HTTP responses; it does not execute page JavaScript. | One page or a small number of pages whose content is present in the returned HTML. |
| Beautiful Soup | Parses fetched HTML or XML so you can select elements and extract text and attributes. | Extracting fields from responses fetched with Requests or another HTTP client. |
| Scrapy | A crawling framework with spiders, scheduling, asynchronous requests, link following, duplicate filtering, exports, pipelines, and crawl controls. | A crawl that needs to visit many pages and be operated, tuned, or exported reliably. |
| Playwright | Controls a real browser from Python, allowing page JavaScript to execute and browser interactions to take place. | Content or navigation that requires rendering, waits, or interaction. |
These tools solve different layers of the problem. Requests downloads; Beautiful Soup parses; Scrapy organizes a crawl; Playwright runs a browser. You can combine them, but do not add browser automation just to parse ordinary HTML.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Before writing a crawler, check whether the site offers a documented API, bulk export, or search endpoint. Those are often more stable and efficient than extracting data from page markup.
#1 Best Overall
Fetch one static page with Requests
A useful first fetch needs more than a URL: validate the scheme, identify your client, set a timeout, handle HTTP errors, and use bounded retries for temporary failures. This example uses https://example.com/, a simple page suitable for checking that the request-and-parse pipeline runs. Change the target only to a site you are permitted to crawl.
from urllib.parse import urlparse
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
def make_session():
retry = Retry(
total=3,
connect=3,
read=3,
status=3,
backoff_factor=1,
status_forcelist=(429, 500, 502, 503, 504),
allowed_methods=frozenset({"GET", "HEAD"}),
respect_retry_after_header=True,
)
adapter = HTTPAdapter(max_retries=retry)
session = requests.Session()
session.mount("http://", adapter)
session.mount("https://", adapter)
session.headers.update({
"User-Agent": "ExampleResearchCrawler/1.0 (contact: [email protected])"
})
return session
def fetch_page(url, session):
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
raise ValueError(f"Expected an absolute HTTP(S) URL, got: {url!r}")
response = session.get(url, timeout=(5, 20), allow_redirects=True)
response.raise_for_status()
return response
if __name__ == "__main__":
with make_session() as session:
response = fetch_page("https://example.com/", session)
print("Requested:", "https://example.com/")
print("Final URL:", response.url)
print("Status:", response.status_code)
print("Content type:", response.headers.get("Content-Type", "not stated"))
print("First 300 characters:")
print(response.text[:300])
What to adapt safely
- Replace the example User-Agent with a descriptive name and a real contact route you control. Do not impersonate a browser or another organization.
- The timeout tuple is connect timeout, then read timeout, in seconds. A timeout prevents a stalled request from blocking a worker indefinitely; tune it for the site and network rather than removing it.
- The retry policy is bounded and limited to GET and HEAD. It honors
Retry-Afterwhere supplied, but retries are not permission to keep hitting a site that is returning rate-limit or server errors. response.urlrecords the final URL after redirects. Keep it with the requested URL: redirects can move a request to another path or host and affect your crawl boundary.
raise_for_status() raises an HTTP error for unsuccessful status codes instead of letting an error page look like valid content. Network errors, timeouts, invalid URLs, and HTTP errors should be logged with the URL and handled by the calling crawl loop.
Parse the response with Beautiful Soup
Parsing is a separate step from downloading. Beautiful Soup receives HTML already in memory and provides tag, attribute, and CSS-selector methods for extracting data. Install the packages with python -m pip install requests beautifulsoup4.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
from bs4 import BeautifulSoup
def extract_page(html, page_url):
soup = BeautifulSoup(html, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else None
heading = soup.select_one("h1")
description = soup.select_one("p")
return {
"url": page_url,
"title": title,
"heading": heading.get_text(" ", strip=True) if heading else None,
"description": description.get_text(" ", strip=True) if description else None,
}
with make_session() as session:
response = fetch_page("https://example.com/", session)
item = extract_page(response.text, response.url)
print(item)
The selectors here are deliberately simple. For a real site, inspect the returned HTML and choose elements tied to the data you need, not incidental layout details. Normalize whitespace with get_text(" ", strip=True), and treat absent elements as missing data rather than assuming every page has the same structure. If the site changes its markup, stable semantic tags or attributes are generally less brittle than deeply nested positional selectors.
Rank #2
Add pagination and crawl boundaries
A crawler needs more than a loop over links. It needs an allowed host boundary, canonicalized URLs, a visited set, a stopping condition, and an explicit pace. The following small queue illustrates those controls. It extracts each page’s title and discovers same-host links; it does not claim that every linked page is relevant, so production crawlers should restrict paths or filter links to the target content.
from collections import deque
from urllib.parse import urldefrag, urljoin, urlparse
import time
from bs4 import BeautifulSoup
def normalized_url(base_url, href):
absolute = urljoin(base_url, href)
absolute, _fragment = urldefrag(absolute)
parsed = urlparse(absolute)
if parsed.scheme not in {"http", "https"}:
return None
return absolute
def crawl(start_url, max_pages=20, delay_seconds=1.0):
start_host = urlparse(start_url).netloc.lower()
queue = deque([(start_url, 0)])
visited = set()
records = []
with make_session() as session:
while queue and len(visited) < max_pages:
url, depth = queue.popleft()
if url in visited:
continue
visited.add(url)
try:
response = fetch_page(url, session)
except requests.RequestException as exc:
print(f"Fetch failed for {url}: {exc}")
continue
# Do not follow a redirect outside the original host.
final_host = urlparse(response.url).netloc.lower()
if final_host != start_host:
print(f"Skipping off-host redirect: {url} -> {response.url}")
continue
soup = BeautifulSoup(response.text, "html.parser")
h1 = soup.select_one("h1")
records.append({
"url": response.url,
"title": soup.title.get_text(" ", strip=True) if soup.title else None,
"heading": h1.get_text(" ", strip=True) if h1 else None,
"depth": depth,
})
if depth == 0:
for link in soup.select("a[href]"):
candidate = normalized_url(response.url, link["href"])
if candidate and urlparse(candidate).netloc.lower() == start_host:
if candidate not in visited:
queue.append((candidate, depth + 1))
time.sleep(delay_seconds)
return records
if __name__ == "__main__":
pages = crawl("https://example.com/", max_pages=20, delay_seconds=1.0)
for page in pages:
print(page)
This deliberately follows links only from the start page, so it demonstrates a bounded one-level crawl rather than an unbounded site spider. A pagination crawler should extract the site’s actual “next” link or page parameter, validate each next URL against the same boundary, and stop when the link is absent or already visited. If a page-number scheme is used, define a maximum page count and stop on an empty or repeated result; never assume pagination continues forever.
Why URL handling matters
urljointurns relative links into absolute URLs. Fragments are removed because they identify positions inside a page, not distinct HTTP resources.- Comparing hosts prevents a crawl from silently expanding to external domains. If subdomains are in scope, specify that rule deliberately instead of treating every hostname as allowed.
- For larger crawls, normalize query parameters only when you understand their meaning. Removing tracking parameters blindly can merge distinct pages; retaining every parameter can create duplicate or infinite URL variants.
- A visited set prevents repeated fetches within the current process. Persist crawl state if a long-running crawl must resume after a restart.
Respect robots.txt and control crawl load
Check the site’s robots.txt and applicable site rules before crawling. Python’s standard-library urllib.robotparser can parse robots.txt and answer whether a user agent may fetch a URL:
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutefrom urllib.robotparser import RobotFileParser
from urllib.parse import urlparse
def allowed_by_robots(page_url, user_agent="ExampleResearchCrawler"):
parsed = urlparse(page_url)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
parser = RobotFileParser()
parser.set_url(robots_url)
parser.read()
return parser.can_fetch(user_agent, page_url)
url = "https://example.com/"
if allowed_by_robots(url):
print("robots.txt permits this URL for the selected user agent")
else:
print("Do not fetch this URL based on the robots.txt result")
This minimal check reads robots.txt but does not add custom timeout, retry, or logging behavior to that read. In a production crawler, fetch and parse the file using the same operational care as other network requests, and decide what to do if it cannot be reached. A robots.txt result is one input, not a complete review of terms, access controls, privacy, or applicable law.
For a Scrapy crawl, its optimization guidance identifies three important controls: CONCURRENT_REQUESTS caps simultaneous downloads, CONCURRENT_REQUESTS_PER_DOMAIN limits simultaneous requests to one domain, and DOWNLOAD_DELAY sets a minimum gap between requests. Read robots.txt, translate any applicable Crawl-delay or Request-rate directives into crawl settings as needed, and increase concurrency gradually. A queue-based script’s sleep is only a basic delay; it does not by itself coordinate multiple workers or enforce a per-domain limit across processes.
Watch for rising 429 or 503 responses, increasing retries, ban or challenge pages, and growing response latency. These are signs to pause or reduce request rates, not to add more workers. Prefer an API, bulk export, or search endpoint when available.
Move to Scrapy when the crawl needs operations
Scrapy describes itself as “an application framework for crawling websites and extracting structured data.” It is the natural next step when the crawl spans many pages, needs asynchronous scheduling, duplicate filtering, feed exports, pipelines, configurable retries or concurrency, or needs to be run repeatedly. Its tutorial demonstrates a spider’s start method, parsing with CSS selectors, following links with response.follow, pagination, and duplicate-request filtering.
Install Scrapy with python -m pip install scrapy. Create a project with scrapy startproject tutorial, then place a spider such as this in tutorial/spiders/example.py:
import scrapy
class ExampleSpider(scrapy.Spider):
name = "example"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/"]
def parse(self, response):
heading = response.css("h1::text").get()
yield {
"url": response.url,
"title": response.css("title::text").get(),
"heading": heading.strip() if heading else None,
}
# Follow same-domain links. Add path and content filters for a real crawl.
for href in response.css("a::attr(href)").getall():
yield response.follow(href, callback=self.parse)
Run it from the project directory and export the extracted items to JSON Lines:
scrapy crawl example -O pages.jsonl
Scrapy handles scheduling and duplicate requests that a hand-built queue would otherwise need to manage. Its overview also describes JSON, CSV, and XML exports, storage backends, selectors, middleware, robots.txt support, and crawl-depth restriction. Configure scope and crawl settings rather than assuming a framework default matches the target site’s rules. For a paginated site, follow only the actual next-page link and let duplicate filtering stop repeated URLs; use a depth restriction when link discovery could fan out beyond the intended section.
The official Scrapy tutorial names Automate the Boring Stuff with Python as a useful resource for new Python programmers. A beginner who is not comfortable with Python functions, loops, exceptions, and dictionaries may find it easier to learn those basics before debugging a spider.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Use Playwright only when a browser is necessary
Requests receives the server response but does not run JavaScript. If the data appears only after browser-side rendering, a click, a dialog choice, or another interaction, Playwright can control a browser from Python and wait for the relevant content. Install the package and browser binaries with python -m pip install playwright followed by python -m playwright install chromium.
Best Value
import asyncio
from playwright.async_api import async_playwright
async def read_rendered_heading(url):
async with async_playwright() as playwright:
browser = await playwright.chromium.launch()
page = await browser.new_page()
try:
response = await page.goto(url, wait_until="domcontentloaded", timeout=30000)
if response is not None and not response.ok:
raise RuntimeError(f"Page returned HTTP {response.status}: {url}")
# Wait for the content this task needs, not an arbitrary long sleep.
heading = page.locator("h1").first
await heading.wait_for(state="visible", timeout=10000)
return {
"url": page.url,
"heading": (await heading.inner_text()).strip(),
"title": await page.title(),
}
finally:
await browser.close()
if __name__ == "__main__":
print(asyncio.run(read_rendered_heading("https://example.com/")))
The sample waits for an h1 to demonstrate a meaningful selector wait; a production target should wait for a locator that identifies the data actually required. domcontentloaded avoids waiting for every resource, while the explicit locator wait covers delayed content. Choose a different load condition only when the page’s behavior requires it. If the data is supplied by a JSON network response, inspect whether that endpoint can be accessed appropriately and reliably rather than automating clicks to read the rendered copy.
When to switch from Requests
- Use Requests plus Beautiful Soup if the data is already present in the fetched HTML.
- Use Scrapy if many HTTP pages need coordinated scheduling, crawl controls, duplicate handling, exports, or repeatable operation.
- Use Playwright when JavaScript execution, browser waits, interactions, or browser-only navigation are necessary.
Browser sessions use more resources than direct HTTP requests, and selectors coupled to a site’s UI can break when the interface changes. Keep browser concurrency conservative, close pages and browsers in cleanup paths, and do not use Playwright to evade access controls or a site’s anti-automation protections.
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than crawl its links and extract structured fields, ScreenshotNeo is a separate API option—not a substitute for a crawler. It returns screenshots or PDFs from one GET request and also provides an MCP server for AI agents. See the ScreenshotNeo website and API documentation.
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server lets AI agents use screenshot, page-information, and PDF-capture tools.
- The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month with no card.
Troubleshoot common crawler failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Timeout or connection exception | Slow response, network issue, or server not accepting the request. | Log the URL and exception, keep a finite timeout, and retry only a bounded number of times. If delays recur, reduce load or stop. |
| HTTP 429 or 503 | The server is rate-limiting requests or is unavailable. | Honor Retry-After where present, lower request rate and concurrency, and reassess whether the crawl should continue. |
| Expected text is missing in parsed HTML | The selector changed, markup differs on this page, or JavaScript adds the content later. | Inspect the fetched response first. If the markup contains the data, fix the selector or missing-field handling; if not, check for an appropriate API or use a browser only if rendering is genuinely required. |
| The same page is fetched repeatedly | URL variants, redirect loops, or missing duplicate tracking. | Record normalized URLs and final redirect URLs, enforce a host/path boundary, and use a visited set or Scrapy’s duplicate filtering. |
| Playwright times out waiting for a locator | The selector is wrong, the content never appears, or it is behind a state your script has not reached. | Check the page URL and locator, wait for the specific content condition, and handle required navigation or dialogs only where permitted. Do not replace the wait with an unlimited sleep. |
| Scrapy discovers too many pages | Broad link following exposes calendars, filters, or query variants. | Restrict allowed domains and paths, filter links, add a depth bound, and track URL variants before expanding the crawl. |
Improve reliability, performance, and output quality
- Keep the first version observable. Log requested URL, final URL, status, elapsed time, and failure reason. Counts of fetched pages, errors, retries, and duplicate skips make a crawl’s behavior visible.
- Bound work. Set page and depth limits while developing. Add a stop condition for empty results and make any pagination logic finite.
- Use structured output. Store records as dictionaries with stable field names, then write JSON Lines or another format suited to downstream use. Preserve source URLs so extracted values can be traced to the page that supplied them.
- Increase throughput cautiously. A synchronous one-worker Requests script is easy to reason about; Scrapy’s asynchronous scheduling is useful for broader workloads. Neither framework makes high request volume acceptable. Respect per-domain limits and watch status codes and latency as concurrency changes.
- Choose the cheapest sufficient layer. HTTP fetches avoid browser startup and rendering overhead. Use Playwright only for pages that require it, and avoid launching a new browser per individual field or request when a controlled browser session can serve the intended work.
- Separate temporary errors from data errors. A failed fetch, absent selector, malformed record, and blocked page need different logs and recovery decisions. Do not silently export error pages as successful data.
Frequently Asked Questions
Does Beautiful Soup download webpages?
No. It parses HTML or XML already obtained by a separate fetch step, such as a Requests call.
Can I use Playwright to crawl an entire website?
It can navigate pages in a browser, but for broad multi-page crawling Scrapy is usually the better scheduling and crawl-management layer. Use browser automation only for pages that need browser behavior.
Should I scrape a site’s API endpoint instead of its page?
Prefer a documented API, export, or search endpoint when one is available and its terms permit your use. It can avoid brittle page selectors, but still apply the site’s access rules and appropriate request limits.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

