What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Web scraping is the automated retrieval of web content and extraction of selected information. A complete workflow may discover pages, fetch their responses, parse the returned documents, extract and validate fields, then save the results. Those are separate jobs: a parser is not a crawler, and a successful HTTP request does not guarantee that its response contains the page data you need.
This guide explains how the pieces fit together, when to choose a parsing library or a crawling framework, and what to consider for JavaScript-rendered pages, robots.txt, safety, and legality.
What is web scraping?
Web scraping is the automated retrieval of web content followed by the extraction of selected information. A small script might request one page and collect its title and price. A recurring job might discover links, visit many pages, extract matching fields, and store validated records.
It helps to separate the workflow into five steps:
- Crawl: find or choose the pages to visit. This can mean following links, reading a sitemap, or using a known list of URLs.
- Fetch: send a request and receive a response. The response may contain HTML, JSON, XML, an error page, or something else.
- Parse: turn the response body into a structure your code can inspect, such as a document tree.
- Extract: select the fields you want, using selectors, traversal, or a documented data interface.
- Validate and output: check that required values exist and have sensible formats before writing them to a file, database, or downstream system.
Keeping these stages distinct makes failures easier to diagnose. If a field is missing, the cause may be a failed request, a changed page structure, content that was never in the response, or an extraction rule that no longer matches.
#1 Best Overall
How are crawling, fetching, parsing, and extraction different?
A crawler manages which pages to visit. A fetcher retrieves a response for a URL. A parser interprets that response’s content. An extractor selects information from the parsed result. Some frameworks coordinate several of these tasks, but the concepts remain different.
For example, Beautiful Soup and lxml are parsing libraries for HTML or XML. Scrapy is an application framework for writing spiders that crawl sites and extract data; it also has CSS and XPath selectors. A Scrapy callback can use Beautiful Soup to parse a response body, so these tools are not mutually exclusive.
For a one-off page, a direct request plus a parser may be enough. A recurring multi-page job is more likely to benefit from a framework’s crawling and extraction structure. That is a practical choice based on the tools’ roles, not a universal speed or performance result.
Recommended Free Tools
Should I use an API or scrape HTML?
Check whether the site provides an official API or structured feed for the data you need. If one is available and its terms and capabilities fit your use, it gives you a defined data interface rather than requiring you to infer fields from page markup. API availability is site-specific; there is no single API that covers every website.
If you fetch a page directly, inspect both the HTTP response and its content before parsing. A response can be received successfully at the network level while still representing an error or containing a different format than expected. In browser JavaScript, a Fetch promise can fulfill for an HTTP error such as 404, so check response.ok or response.status before treating the body as usable. Then choose the right body reader, such as text or JSON, for the response format.
What should I use: Scrapy, Beautiful Soup, or lxml?
| Tool | Role | When it may fit |
|---|---|---|
| Scrapy | A framework for spiders, crawling, and extraction; includes CSS and XPath selectors. | A recurring or multi-page crawl that needs a framework to organize visits and extraction. |
| Beautiful Soup | A library for parsing HTML and XML. | A script or application that needs to inspect and navigate a parsed document; it can also be used inside Scrapy callbacks. |
| lxml | A library for parsing HTML and XML. | A task that calls for a parsing library, including when working with document structure and selectors. |
The key decision is not simply which name is most popular. Ask whether you need to coordinate page discovery and repeated requests, or only parse a response you already have. You can combine a crawling framework with a parser when that division suits the task.
How do I parse a page and validate extracted data?
Start with one response and a small set of required fields. Inspect the actual response body, choose a parser appropriate to its format, and build extraction rules against the structure that is present—not the structure you assume the site returns. CSS selectors, XPath, and parser-specific traversal are common approaches; the best fit depends on the document and the tool.
Validation belongs after extraction. Check for absent values, unexpected types, malformed dates or URLs, and fields that are present but empty. Decide explicitly what should happen when a required value is missing: reject the record, log it for review, or store a clearly represented null. Do not silently turn a missing value into a plausible-looking default.
Real markup may be malformed or change over time. Choose a parser that can handle the source format, and treat the page structure as an input that can evolve. If the site changes a class name or nests a field differently, extraction can fail even when fetching still works. Logs that distinguish fetch failures from parse and validation failures make this easier to spot.
How do I handle JavaScript-heavy pages?
First find out whether the needed data is already in the initial HTML response or is retrieved separately by page JavaScript. Browser code can fetch network resources that return JSON, HTML, or text; the visible page may therefore be assembled from additional requests rather than being fully represented in the first response.
Rank #3
If the data is available in an authorized structured response, use an appropriate request and parse that format. If the content only becomes available after browser rendering, use a rendering approach suitable for the site’s access rules. There is no one browser automation method established for every site, and a screenshot is not a substitute for structured extraction when you need machine-readable fields.
For a visual capture rather than an extraction pipeline, ScreenshotNeo is a website screenshot API and MCP server. Its 63 options include full-page capture, element selection, waits, custom CSS and JavaScript, and PDF output. It can help capture a rendered page, but it should not be mistaken for a general-purpose crawler or parser.
What is robots.txt, and does it give permission?
robots.txt is a text file through which a site publishes crawler rules. The Internet Engineering Task Force’s RFC 9309, published in September 2022, defines how crawlers interpret rules and groups. It also states: “These rules are not a form of access authorization.” In other words, robots.txt is a crawler instruction protocol, not a login, access-control mechanism, or grant of legal permission.
Crawlers should follow applicable parseable rules. Google’s documentation says Google crawlers download and parse robots.txt before crawling. For a Python check, urllib.robotparser.RobotFileParser reads and parses robots.txt, and can_fetch(useragent, url) reports whether that user agent is allowed to fetch the URL under the parsed rules. The documentation page surfaced for Python 3.16.0a0, a prerelease; verify behavior against the Python version you deploy.
Respecting robots.txt is only one part of deciding whether and how to collect data. It does not replace checking terms, access controls, privacy obligations, and applicable law.
Is web scraping legal?
There is no reliable universal yes-or-no answer. The legal outcome depends on the conduct, the data, the site, and the relevant jurisdiction. A Cornell Legal Information Institute explainer focused on US law discusses a Ninth Circuit decision in which accessing publicly available data was not treated as access without authorization under the Computer Fraud and Abuse Act in that case. The same explainer describes limits around circumventing protective measures.
That summary does not resolve every site’s terms, every kind of collection, or laws outside the circumstances it discusses. Contract claims, privacy and data-protection rules, intellectual property, access controls, and local law may all matter. If the stakes are material, get advice for the actual facts and jurisdiction instead of treating a general article as a legal determination.
How can I avoid overloading a website?
- Follow the site’s published crawler instructions and avoid fetching pages you do not need.
- Use a conservative request pace and reduce it further if the site signals overload or denies access.
- Stop or back off when requests fail in ways that indicate the site cannot or will not serve them.
- Keep track of requests and failures so a retry loop does not repeatedly hit a struggling server.
RFC 9309 defines robots rules and their handling; it does not set one universally safe request rate for every site. Choose a pace appropriate to the site and your use, and do not interpret an allowed robots.txt path as permission to send unlimited requests.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should I handle unsafe or unusually large content?
Fetched markup should be treated as untrusted input. Parsing HTML into a separate document does not make it safe to move nodes into your live browser page: inserting unsafe content can create a cross-site scripting risk. Sanitize content or use Trusted Types before insertion, and avoid injecting scraped markup as active page content unless you have handled that risk.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Large responses also need resource limits. Scrapy’s security documentation discusses response-size and parser limits, which can protect resources but may truncate unusually large content. Set limits with the expected page sizes in mind, and treat truncation as a possible cause when a response looks incomplete rather than assuming the source omitted data.
Best Value
Or skip the browser setup
If the task is to capture a screenshot or PDF of a rendered page rather than build your own browser-rendering workflow, ScreenshotNeo accepts one GET request with a URL and returns an image or PDF. Example using cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options. Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides screenshot tools for AI agents. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
What should I check before choosing an approach?
- Scope: Is this one page, a supplied URL list, or a recurring multi-page crawl?
- Format: Does the response contain HTML, XML, JSON, or another format?
- Page behavior: Is the data in the initial response, or fetched by client-side code?
- Extraction: Is a documented API available, or will CSS, XPath, or parser traversal work?
- Operations: Do you need retries, request pacing, logs, deduplication, and field validation?
- Compliance and safety: Have you considered crawler rules, terms, access controls, privacy, and the relevant jurisdiction?
Frequently Asked Questions
Can I use Scrapy with Beautiful Soup?
Yes. Scrapy’s callbacks receive responses that can be parsed with Beautiful Soup; using both does not change their distinct roles.
Does a robots.txt rule mean a page is public?
No. Robots.txt communicates crawler rules; it is not an access-control system or legal authorization.
Will a successful request always contain the page I see in my browser?
No. The response may be an error, a different format, or content that lacks data later loaded by client-side code.
Does parsing HTML make it safe to display?
No. Treat fetched markup as untrusted, especially before inserting parsed nodes into an active browser document.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

