Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsMost scraping problems are solved by identifying what the site actually sends, choosing the least complicated permitted way to retrieve it, and checking that the extracted data is still correct. Start with the page’s network requests before reaching for a browser; slow down and respect access rules; treat 403s and CAPTCHA pages as signals to stop or use an approved route, not obstacles to defeat. The right approach depends on whether you need structured data, browser-rendered content, or simply a visual capture.
Start by identifying what is failing
A scraper can fail at several different stages: it may not be allowed to access a page, it may fetch a page without the data you need, it may extract the wrong fields, or it may collect data that you should not use. Diagnose those stages separately. Save the response status, final URL, relevant headers, a short response-body sample, and extraction results for each request. That evidence helps distinguish an access restriction from a JavaScript-rendering issue or a broken selector.
- Access: Did the request succeed, and did the site return the page you expected rather than a redirect or challenge?
- Availability: Is the desired information present in the returned HTML or in a data request the page makes?
- Extraction: Did your parser find the expected fields, and are their values plausible?
- Permission and use: Are the collection method and intended use consistent with the site’s rules and applicable law?
Keep a small set of representative URLs as regression checks. A successful HTTP response is not proof that a scrape is complete: a page can load while a selector returns nothing, or while a challenge page occupies the response body.
When the page returns 403, a CAPTCHA, or a challenge
A 403 means the server refused the request; it does not, by itself, reveal the reason. A site may apply web application firewall rules, IP controls, JavaScript-based checks, authentication requirements, geographic rules, or rate limits. A CAPTCHA or challenge page is especially clear evidence that the site is restricting automated access.
#1 Best Overall
- Check that you are requesting the correct public URL and that the site has not redirected you to a sign-in page, error page, or challenge.
- Review the site’s terms and published access guidance, and confirm you have permission for the collection and intended use.
- If you have permission, reduce your request rate and concurrency, avoid unnecessary repeat requests, and cache responses where appropriate.
- Look for an official API, data export, licensed feed, or other approved access route. Ask the site owner for access if none is documented.
- If access remains restricted, stop. Do not attempt to bypass a CAPTCHA, WAF, authentication boundary, or other protective measure.
Repeated retries are not a fix for an access denial. They can increase load and make a temporary block worse. Record the status and response pattern, then handle the page as a restricted or unavailable result rather than feeding it into the normal extraction pipeline.
When JavaScript hides the data
First inspect the page’s network activity in your browser’s developer tools. Look for requests that return the data as JSON or another structured response. Scrapy’s guidance for dynamic pages similarly recommends finding and reproducing the underlying data request when possible; a browser-rendered DOM is needed only when the content depends on browser behavior that cannot reasonably be retrieved that way.
Prefer the underlying data request when it is permitted
If the browser requests a JSON endpoint that contains the fields you need, inspect its URL, method, parameters, and response. Confirm that the endpoint is intended for use under the site’s rules; a request being visible in a browser does not automatically grant permission to automate it. Reproducing a permitted data request is often simpler than loading an entire browser, and gives you structured fields instead of markup to parse.
Use browser automation when rendering is necessary
Use a headless browser such as Playwright when the needed content appears only after browser-side interactions or rendering and there is no suitable approved data route. Wait for a meaningful condition, such as a known content element, rather than relying on an arbitrary long sleep. Keep the same access, pacing, and permission safeguards as with direct HTTP requests. A browser can render a page; it does not make restricted access permissible.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →For a browser-rendered page, save enough context to diagnose a failure: the final URL, whether the target element appeared, and whether the page instead showed a sign-in, error, or challenge screen. Avoid treating an empty DOM query as an empty dataset until you have checked that the page finished rendering the content.
Choose the simplest tool that fits the job
There is no universally best scraper. Pick based on the source’s permitted access method, the data’s availability, volume, and the effort you can spend on monitoring and maintenance.
Rank #3
| Approach | Best fit | Main trade-off |
|---|---|---|
| Direct HTTP client | Public, server-delivered HTML or a permitted structured endpoint | Lightweight and direct, but does not execute browser JavaScript |
| Scrapy | A crawl with multiple pages, structured extraction, and configurable request pacing | Requires parser maintenance and explicit configuration of crawl pacing directives |
| Playwright or another browser automation framework | Content that genuinely depends on browser rendering or interaction | More resource-intensive and operationally involved than fetching the underlying data directly |
| Managed scraping service | A team that wants an externally managed capture or scraping workflow | Evaluate cost, observability, data quality controls, authentication handling, and whether the service’s approach complies with the site’s rules |
For managed scraping services, compare them on completeness, throughput, latency, infrastructure cost, maintenance, observability, data quality controls, authentication handling, and compliance posture. Ryan Mitchell’s Web Scraping with Python, 3rd Edition discusses JavaScript-heavy pages, blocking, legal considerations, and managed-service resources; its publisher lists ScrapingBee, ScraperAPI, and Zyte among those resources. That is a starting point for evaluation, not an endorsement or a statement about their current capabilities.
Respect robots.txt and control crawl load
Read a site’s robots.txt instructions before crawling and follow applicable site policies. Robots.txt is a crawl instruction, not a universal legal prohibition or a mechanism that hides pages from all visitors. Google Search Central specifically cautions against using it to hide pages from search results.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Scrapy can use its robots middleware, but the framework’s optimization guidance says it does not automatically enforce Crawl-delay or Request-rate directives. Translate applicable pacing instructions into explicit delay and concurrency settings yourself. For example, a cautious Scrapy configuration could start like this, with the values adjusted to the site’s published policy and your permission:
ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 1
Those example values are not a universal safe rate or a substitute for the target’s instructions. Use fewer simultaneous requests and longer delays when the site asks for them or when your crawl causes load. Cache responses, deduplicate URLs, and avoid fetching unchanged pages more often than necessary. On transient server failures, use bounded retries with increasing delays; do not retry access denials or challenge pages as though they were ordinary network glitches.
Keep extraction accurate as pages change
Websites change markup, labels, and page structure. A selector can keep running while quietly returning incomplete or incorrect data. Separate fetching from parsing and validation so that a successful request cannot silently become a successful-looking bad record.
- Normalize fields: Convert values into consistent formats and handle missing optional values explicitly.
- Validate required data: Check that key fields exist and meet basic expectations before accepting a record.
- Detect duplicates: Use a stable identifier where available, and log records that collide unexpectedly.
- Track extraction health: Record response codes, selector matches, missing-field counts, and record counts by page or run.
- Alert on drift: Treat sudden drops in extracted records or increases in missing fields as a possible layout or access change.
- Test parser changes: Keep representative saved responses or fixtures and run them against the parser when selectors change.
Version parsers and preserve enough run metadata to compare results over time. When the site changes, pause publication or downstream updates until you know whether the change is a real data change, a new layout, or a response that should never have been parsed as content.
Recommended Free Tools
Best Value
Is web scraping legal?
There is no single answer for every site, jurisdiction, and use. Cornell Law School’s Legal Information Institute summarizes that screen scraping is technically legal in general, while noting that circumventing typical protective measures can create exposure under the U.S. Computer Fraud and Abuse Act. That broad summary is not a guarantee that a particular scrape is lawful.
Before collecting or republishing data, consider the site’s terms, whether access requires authentication, any access controls, copyright, privacy obligations, the data subjects and intended use, and the laws that apply where you and the site operate. Public availability alone does not settle those questions. If the data is sensitive, the use is commercial, or access terms are unclear, get legal advice or obtain permission rather than relying on a general rule or a past court case as a blanket safe harbor.
Troubleshoot common scraper failures
| Symptom | Likely explanation | What to do |
|---|---|---|
| 403 or repeated denial | Access policy, permissions, IP controls, or another restriction | Check permission and published routes, reduce load if authorized, use an approved API, or stop. Do not evade the restriction. |
| CAPTCHA or challenge HTML | The site is asking to verify or restrict automated access | Do not solve or bypass it with automation; seek an approved access method or permission. |
| HTML loads but fields are absent | Data may be injected by JavaScript, retrieved separately, or the markup may have changed | Inspect browser network requests and verify the response body before changing selectors; use browser rendering only if needed. |
| Some pages work and others fail | Different page templates, authentication states, geography, or access rules | Compare final URLs, response codes, and page templates; handle each permitted case explicitly. |
| Records suddenly disappear or duplicate | Selector drift, changed identifiers, pagination behavior, or a partial response | Check validation metrics and saved examples; pause downstream updates until the parser and source response are understood. |
| Requests slow down or begin failing under load | Too much concurrency, repeated fetching, transient server trouble, or a site limit | Lower concurrency, increase delays, cache and deduplicate, then use bounded backoff for transient failures. Respect denials rather than retrying them. |
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than extract structured records, ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. For a visual capture of Stripe’s homepage in WebP:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.
Frequently Asked Questions
Does robots.txt prevent a page from appearing in Google Search?
No. Google Search Central says robots.txt should not be used to hide pages from search results; it is not a substitute for access control.
Is a screenshot API a replacement for a data scraper?
Not when you need structured fields or records. A screenshot API returns a visual capture, while scraping for data requires retrieving and validating the underlying information.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

