Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
How do I scrape a web page with Python? Start with a small, permitted page and a five-step loop: request the URL, check the response, parse its HTML, select the fields you need, and save validated records. The example below uses Requests for HTTP and Beautiful Soup for parsing. Once the page is static and the selectors are understood, you can move to Scrapy for repeatable crawls or Playwright when content truly requires a browser.
The web-scraping loop
Web scraping is not one operation. An HTTP client downloads a response; a parser turns the response body into a document; selectors locate elements; extraction converts them into values; and validation and storage produce usable data.
- Request: send an HTTP GET (or use an authorized API when one exists).
- Inspect: check the status code, content type, and a small part of the returned HTML.
- Parse: build a searchable HTML tree.
- Select and extract: use narrow CSS or XPath selectors, then read text or attributes.
- Clean, validate, and save: normalize whitespace, handle missing fields, check sample records, and write CSV or JSON.
A page that looks complete in a browser may return only a shell before JavaScript runs. Test the initial response before choosing a browser tool.
Free tools Windows power users keep installed
One-click scans. No signup required.
A complete static-page example
Use a page intended for practice, such as the tutorial site in the Scrapy tutorial, and replace the selectors after inspecting its current markup. Install the libraries in an isolated environment:
#1 Best Overall
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4
The following script requests a page, extracts a title and repeated records, and writes JSON. The selectors are examples; inspect your target page and change them rather than assuming every site has the same structure.
from __future__ import annotations
import json
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
URL = "https://quotes.toscrape.com/"
headers = {"User-Agent": "intro-scraper/1.0 (contact: [email protected])"}
response = requests.get(URL, headers=headers, timeout=30)
response.raise_for_status()
content_type = response.headers.get("content-type", "")
if "html" not in content_type.lower():
raise ValueError(f"Expected HTML, got {content_type!r}")
soup = BeautifulSoup(response.text, "html.parser")
page_title = soup.title.get_text(" ", strip=True) if soup.title else ""
records = []
for card in soup.select(".quote"):
text_node = card.select_one(".text")
author_node = card.select_one(".author")
tag_nodes = card.select(".tags .tag")
if text_node is None or author_node is None:
continue
records.append({
"text": text_node.get_text(" ", strip=True),
"author": author_node.get_text(" ", strip=True),
"tags": [tag.get_text(" ", strip=True) for tag in tag_nodes],
"author_url": (urljoin(URL, author_node.get("href"))
if author_node.get("href") else None),
})
if not records:
raise RuntimeError("No records found; inspect the HTML and selectors")
output = {"url": URL, "title": page_title, "records": records}
with open("quotes.json", "w", encoding="utf-8") as file:
json.dump(output, file, ensure_ascii=False, indent=2)
print(f"Saved {len(records)} records")
raise_for_status() turns 4xx and 5xx responses into exceptions instead of silently parsing an error page. The content-type check catches redirects to a login page, an API response, or a download. get_text(" ", strip=True) collapses line breaks and surrounding spaces. get("href") reads an attribute; it is different from extracting visible text.
Inspect before you extract
When a selector returns nothing, save or print a short response sample and inspect the source HTML, not only the rendered browser view:
print(response.status_code, response.url)
print(response.text[:1000])
print(soup.select_one(".quote"))
Start with semantic containers such as an article, product, row, or card, then select fields inside that container. This prevents a page-wide selector from mixing navigation, adverts, and content.
Rank #2
Missing fields and validation
Real pages omit fields. Test nodes before calling methods, choose an explicit default, and reject records that lack an identifying value. Validate a handful of rows manually before collecting hundreds. Keep the source URL and retrieval time with each record when later auditing matters.
CSS selectors and XPath
Beautiful Soup accepts CSS selectors through select(): article.product h2 means an h2 inside a product article; [data-id] selects elements carrying that attribute. Prefer stable classes, IDs, semantic tags, or documented data attributes over generated class names.
XPath is useful when you need traversal or predicates, such as “the link whose text is Next” or an element following a particular heading. Scrapy selectors support both CSS and XPath and are built over Parsel, which uses lxml. The Scrapy selector guide documents the syntax. Beautiful Soup is forgiving of imperfect markup, while the same guide notes a speed drawback; do not treat that note as a universal benchmark for your workload.
Following pagination safely
For a small script, follow a next-page link until it is absent, a maximum page count is reached, or a duplicate URL appears:
from urllib.parse import urljoin
seen = set()
url = "https://quotes.toscrape.com/"
all_records = []
for _ in range(20):
if url in seen:
break
seen.add(url)
page = requests.get(url, headers=headers, timeout=30)
page.raise_for_status()
soup = BeautifulSoup(page.text, "html.parser")
for card in soup.select(".quote"):
text = card.select_one(".text")
author = card.select_one(".author")
if text and author:
all_records.append({"text": text.get_text(" ", strip=True),
"author": author.get_text(" ", strip=True)})
next_link = soup.select_one("li.next a[href]")
if not next_link:
break
url = urljoin(page.url, next_link["href"])
print(f"Collected {len(all_records)} records from {len(seen)} pages")
Use page.url when resolving relative links because redirects can change the base URL. Add a delay between requests, and stop when the site signals an error rather than retrying indefinitely.
Choosing the right Python tool
| Situation | Starting choice | Reason |
|---|---|---|
| A few pages; data is in the returned HTML | Requests plus Beautiful Soup or lxml | Small, understandable separation between retrieval and parsing. |
| Many pages, pagination, link following, repeatable jobs, exports | Scrapy | Project and spider workflow, scheduling, selectors, feed exports, and crawl controls. |
| Data appears only after JavaScript or interaction | Playwright for Python | Automates a browser and exposes request, response, redirect, and resource information. |
| An official API supplies the records | Use the API, subject to its terms | It is generally less fragile and adds less page load than scraping HTML. |
When Scrapy is the next step
Create a project, define a spider, yield dictionaries or items from parse, follow the next link, and export a feed. The official tutorial walks through that workflow. Scrapy selectors support CSS and XPath, and its scheduler handles many requests without you writing the pagination loop yourself.
When Playwright is justified
Use a browser only when the needed data is absent from the initial response or requires permitted interaction such as clicking a tab. Playwright’s Python Request API documents how to observe network requests and responses. Before launching a browser, look for an authorized API or data feed visible in the page’s normal network activity.
Respectful and maintainable crawling
- Identify the crawler with a descriptive User-Agent and a contact address. The Scrapy tutorial explains that owners can then ask you to adjust the crawler rather than block it.
- Read current site instructions, terms, and any API documentation. A
robots.txtfile is not legal advice or proof that a use is permitted. - Keep scope narrow: collect only needed fields, restrict allowed domains and paths, and cache responses when appropriate.
- Control load with delays, per-domain concurrency limits, and AutoThrottle where your framework supports them. Scrapy documents these controls in its overview.
- If using Scrapy, enable RobotsTxtMiddleware and set
ROBOTSTXT_OBEY = Truewhen that matches your project policy. Its middleware documentation describes the configuration and parser behavior. A basic Requests script does not automatically obey robots.txt. - Stop when access is denied or the operator objects. Do not bypass CAPTCHAs, authentication, rate limits, or other access controls.
Permission, privacy and data-protection duties, copyright or database rights, and other rules depend on your jurisdiction, the data, and your access method. Obtain permission or use a supported API when those facts are unclear.
Or skip the browser setup
If your goal is simply a clean image or PDF of a page, ScreenshotNeo makes one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status.
It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every plan includes the features below; the free tier is 1,000 screenshots per month without a card, and paid plans start at $5 for 3,000 shots.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the parameter reference and authentication details in the ScreenshotNeo documentation. Sign up for 1,000 free screenshots a month with no card.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Troubleshooting common failures
403, 429, or a block page
The server may require authorization, a slower rate, or a supported API. Confirm terms and credentials, reduce concurrency, identify your crawler, and stop rather than attempting evasion.
HTML contains no expected records
You may have received a login, consent, error, or JavaScript shell. Print status, final URL, content type, and a response excerpt. Check whether an API supplies the records; use Playwright only if browser execution is genuinely required.
Best Value
Selectors broke after a redesign
Inspect current markup, replace brittle generated classes with stable attributes or semantic containers, and add a validation check that alerts you when record counts unexpectedly fall.
Timeouts and partial output
Set explicit connect and read timeouts, save progress incrementally, retry a small number of transient failures with backoff, and record failed URLs for review. Do not retry permission errors forever.
Encoding or whitespace looks wrong
Use the response’s declared encoding when possible, preserve Unicode with UTF-8 output, and normalize text only after deciding whether line breaks or non-breaking spaces carry meaning.
FAQ
Frequently Asked Questions
How do I extract data from a website using Python?
Use an HTTP client such as Requests to obtain the page, parse the HTML with Beautiful Soup or lxml, select the fields, clean and validate values, then save JSON or CSV. If the provider offers an authorized API, use it instead.
Should I use Beautiful Soup, Scrapy, or Playwright?
Choose Beautiful Soup for a small static task, Scrapy for repeatable multi-page crawls and exports, and Playwright only when permitted browser behavior is needed. Base the choice on initial HTML, page count, interaction, maintenance, and crawl controls.
Does robots.txt make scraping legal?
No. It is an instruction mechanism, not legal advice or proof of permission. Check terms, authorization, privacy obligations, intellectual-property rules, and applicable law; stop if the operator objects.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

