The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Web scraping is the programmatic collection of information from websites, followed by cleaning and storing that information in a usable format such as JSON, XML, CSV or a database. A typical scraper discovers URLs, sends HTTP requests, parses HTML or browser-rendered pages, extracts selected fields, follows links or pagination, validates the results and records where each value came from.
The difficult part is not downloading one page. Reliable scraping requires choosing an authorized access method, controlling request volume, handling JavaScript and failures, respecting robots.txt and site terms, and monitoring selectors when a site changes. This guide explains the complete workflow and includes a runnable Python example.
What web scraping includes
The National Network of Libraries of Medicine describes web scraping as programmatically and systematically collecting information on the web and processing it into analyzable formats that can be serialized, such as JSON or XML, and stored for later use. In practical terms, a scraper turns pages designed for people into records designed for analysis.
- Fetching: requesting an HTML document, JSON response, feed or rendered page.
- Parsing: reading the response and locating elements with CSS selectors, XPath or a structured-data parser.
- Structuring: mapping values into named fields such as title, price, date and URL.
- Quality control: normalizing text, dates and currencies; removing duplicates; checking required fields.
- Storage: writing records to JSON, CSV, an object store or a database with source and retrieval metadata.
Scraping is different from copying a page once by hand. It is repeatable, scheduled and explicit about the fields and pages it collects.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
How a web-scraping workflow works
1. Define the dataset and permission
Write down the fields you need, the URL scope, how often data should be refreshed, how long it will be retained and who may use it. Check whether the publisher offers an API, feed or downloadable dataset; that channel is normally more stable and easier to govern than parsing presentation HTML. Decide in advance what you will do with personal data, credentials and records that are later removed.
2. Discover URLs
Seed the job with known pages, a sitemap, a feed, search results or links exposed by the site. Keep an allowlist of domains and URL patterns. Canonicalize URLs and remove tracking parameters when they do not identify a distinct resource, otherwise the same page can be downloaded repeatedly.
3. Schedule requests
A crawler maintains a queue, prioritizes URLs, limits concurrency and avoids duplicate work. Scrapy calls this component the scheduler. A production queue also records retry counts, next-attempt times and the reason a URL was deferred.
4. Download responses
The downloader handles HTTP headers, cookies, compression, redirects, timeouts, retries and connection limits. Identify your client honestly where appropriate, use a bounded timeout and treat 429, 403 and repeated 5xx responses as signals to slow down or stop rather than as invitations to increase traffic.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches5. Parse and select fields
Use an HTTP parser when the required data is in the response HTML or embedded structured data. Select fields with stable CSS or XPath expressions and record the selector version used. If JavaScript creates the content only after load, use an authorized browser-rendering step or look for an official endpoint that supplies the data.
6. Follow pagination and links
Extract the next-page URL or an API cursor, enqueue it, and stop at a defined page limit or when no next link remains. Keep a visited set keyed by canonical URL. For link discovery, restrict following to the intended domain and path patterns.
7. Normalize and validate
Strip incidental whitespace, parse dates with an explicit timezone, convert currencies only when the source currency is known, normalize Unicode and distinguish an absent value from an empty string. Validate required fields and reject or quarantine records that fail schema checks. Preserve the original value when a transformation could lose information.
8. Store or export
Write items through a pipeline to JSON, CSV, S3-compatible object storage or a database. Store the source URL, retrieval timestamp, HTTP status, parser version and relevant transformation details alongside the extracted fields. This provenance makes an individual value auditable.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match9. Monitor drift
Track status codes, response sizes, field completeness, duplicate rates, selector failures and crawl volume. Alert when a normally populated field becomes empty or a site redesign changes the HTML. A successful process that silently produces empty records is a failed scraper.
Crawling versus scraping
| Activity | Primary purpose | Typical output |
|---|---|---|
| Crawling | Discovering, queuing and fetching pages or links | Responses, URLs and crawl metadata |
| Scraping | Selecting and structuring information from responses | Validated records such as JSON rows |
One program often does both: a crawler schedules requests while a spider parses each response into items. The distinction is useful because URL discovery and field extraction have different failure modes and controls.
Choose the least complex access method
Official API, feed or export
Prefer an official interface when it supplies the fields you need. It usually has a documented schema, clearer rate limits and less selector maintenance. Confirm authentication, pagination, retention and licensing terms.
Static HTML or embedded data
Direct HTTP parsing is efficient when the server response already contains the content. Inspect the response, not only what a browser displays, and check for JSON-LD or other structured data before writing fragile selectors.
JavaScript-rendered pages
Use browser automation or rendering only when the required data is absent from the initial response. Rendering costs more CPU and time and introduces browser, cookie and consent states. Do not use it to bypass a login, paywall or anti-bot control without explicit authorization.
Authenticated or restricted areas
Treat credentials, paywalls and bot checks as access boundaries. Obtain permission, use the provider’s supported API or export, protect secrets and define a stop condition if the operator blocks the account.
Rank #3
Scrape a simple page with Python
This example requests a page, extracts article cards, follows a small number of pagination links and writes records with retrieval timestamps. It assumes the page is publicly accessible and that its terms permit the activity.
Install dependencies
python -m pip install requests beautifulsoup4
Complete script
from datetime import datetime, timezone
from urllib.parse import urljoin
import json
import time
import requests
from bs4 import BeautifulSoup
START_URL = 'https://example.com/news'
MAX_PAGES = 5
HEADERS = {'User-Agent': 'ExampleResearchBot/1.0 ([email protected])'}
session = requests.Session()
session.headers.update(HEADERS)
records = []
seen = set()
url = START_URL
for _ in range(MAX_PAGES):
if not url or url in seen:
break
seen.add(url)
response = session.get(url, timeout=30)
if response.status_code == 429:
raise RuntimeError('The site is rate limiting the crawler; stop and retry later.')
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
retrieved_at = datetime.now(timezone.utc).isoformat()
for card in soup.select('article.card'):
link = card.select_one('a')
title = card.select_one('.title')
if not link or not title:
continue
records.append({
'title': title.get_text(' ', strip=True),
'url': urljoin(response.url, link.get('href', '')),
'retrieved_at': retrieved_at,
'source_page': response.url,
})
next_link = soup.select_one('a[rel="next"]')
url = urljoin(response.url, next_link['href']) if next_link else None
time.sleep(1.0) # keep request rate deliberately low
with open('records.json', 'w', encoding='utf-8') as output:
json.dump(records, output, ensure_ascii=False, indent=2)
Replace the selectors with selectors verified against the target site’s HTML. Keep source_page and retrieved_at; they are essential when a value later needs to be checked. For a real job, add schema validation, persistent retries with backoff, caching and tests that fail when required selectors disappear.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →When to use Scrapy
Scrapy is a Python framework for larger asynchronous crawls. Its architecture centers on an engine, scheduler, downloader, spiders, items, pipelines and feed exports. It provides CSS/XPath selectors, concurrency controls, retries and robots.txt middleware. The official site listed version 2.19.0 in September 2026; check the current release information before pinning a dependency.
import scrapy
class ArticleSpider(scrapy.Spider):
name = 'articles'
start_urls = ['https://example.com/news']
custom_settings = {'ROBOTSTXT_OBEY': True}
def parse(self, response):
for card in response.css('article.card'):
yield {
'title': card.css('.title::text').get(default='').strip(),
'url': response.urljoin(card.css('a::attr(href)').get()),
}
next_url = response.css('a[rel="next"]::attr(href)').get()
if next_url:
yield response.follow(next_url, callback=self.parse)
Run a spider with a feed export such as scrapy crawl articles -O records.json. Set an explicit concurrency and delay instead of accepting aggressive defaults.
What robots.txt does—and does not do
Google Search Central defines robots.txt as a file that tells crawlers which URLs they may access and primarily uses it to manage crawler traffic. Digital.gov likewise describes it as instructions for internet bots that can help manage site performance. It is not an access-control system and does not remove a URL from search results by itself.
Read the file at the site’s root, apply the rules for your user-agent and honor applicable Disallow and Crawl-delay guidance. Scrapy’s ROBOTSTXT_OBEY setting enables its middleware. Scrapy also documents an explicit per-request override that ignores the file; using that is a deliberate governance decision requiring a documented reason and authorization, not a normal fallback.
Is web scraping legal?
There is no worldwide yes-or-no rule. Exposure depends on jurisdiction, authorization, the site’s terms, the kind of data, privacy and copyright interests, database rights, technical behavior and the purpose of the collection. Public visibility is not a universal permission to copy, republish or create profiles.
In its April 18, 2022 preliminary-injunction opinion in hiQ Labs v. LinkedIn, the U.S. Court of Appeals for the Ninth Circuit analyzed whether the Computer Fraud and Abuse Act’s “without authorization” language reaches publicly viewable LinkedIn profiles. The opinion also noted that other claims—such as copyright, breach of contract, trespass to chattels, unjust enrichment, conversion or privacy claims—may still be available. That decision is not a blanket license to scrape every public site, and it does not settle laws outside the circumstances and jurisdiction considered.
- Review terms of use, API documentation, robots.txt and licensing instructions.
- Obtain permission or use an official channel for sensitive, restricted or commercially important data.
- Collect only fields necessary for a defined purpose and protect personal information and credentials.
- Honor rate limits, stop on explicit operator contact and keep an audit trail of decisions.
Operate a scraper responsibly at scale
Rate, concurrency and caching
Use the lowest request rate and concurrency that meets the requirement. Cache successful responses, avoid duplicate URLs and schedule refreshes according to how often the source changes. Exponential backoff for transient errors is safer than immediate retries.
Data quality
Validate schemas, deduplicate by a stable key, retain raw values where practical and record every transformation. Separate parsing errors from genuinely missing source data so a redesign does not appear as a legitimate run of nulls.
Free tools Windows power users keep installed
One-click scans. No signup required.
Security and privacy
Keep API keys and session cookies out of source control and logs. Encrypt stored credentials, restrict access to raw personal data and set deletion periods. Do not collect passwords or defeat authentication controls.
Change management
Version selectors and parser code, maintain fixture pages for tests and alert on field-completeness or response-shape changes. A manual review queue is useful when a high-value field suddenly fails validation.
Common failures and fixes
- 403 or a bot-check page: stop rather than trying to evade it; confirm permission and use an official API or contact the operator.
- 429 responses: reduce concurrency, honor the stated limit, cache results and retry later with backoff.
- Empty selectors: inspect the raw response. The content may be JavaScript-rendered, the selector may have changed or a consent page may have been returned.
- Timeouts: set a finite connect/read timeout, reduce concurrency, retry only transient failures and record the URL for review.
- Duplicate records: canonicalize URLs, persist a visited set and deduplicate on a source identifier plus normalized fields.
- Broken characters or dates: respect the response encoding, normalize Unicode and parse dates with an explicit locale and timezone.
- Pagination loops: track visited URLs and enforce a maximum page count or item count.
- Silent data loss after a redesign: add required-field assertions and alerts on completeness, not just HTTP success.
When a screenshot service helps a scraping workflow
A screenshot service is not a replacement for an API or an HTML parser when you need structured records. It is useful when you need a visual snapshot of a rendered page for QA, archival evidence, human review or an AI agent, especially when browser setup is the expensive part of the pipeline.
Or skip the browser setup:
ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response reports the result with X-Page-Verdict and X-Billed headers.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options. The same service supports PNG, JPEG, WebP and PDF output; full-page capture with lazy images, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, hidden selectors, selector/delay/network-idle waits, blocking ads, trackers, requests or resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
Best Value
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can request captures without you building browser orchestration.
| Plan | Included shots per month | Price |
|---|---|---|
| Free | 1,000 | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
FAQ
Does scraping require a browser?
No. If the response contains the data, an HTTP client and parser are simpler and faster. A browser is needed only for client-rendered content or interactions that an authorized workflow must perform.
Recommended Free Tools
What should I keep so another person can reproduce a record?
Keep the source URL, retrieval time, response status, raw response or an immutable snapshot where permitted, parser and selector versions, and a log of transformations.
How often should a scraper run?
Match the schedule to the source’s change rate and your business need. A slower schedule reduces load and cost; increase it only when fresher data has a defined value and the operator’s limits allow it.
Frequently Asked Questions
Does scraping require a browser?
No. If the response contains the data, an HTTP client and parser are simpler and faster. A browser is needed only for client-rendered content or authorized interactions.
What should I keep so another person can reproduce a record?
Keep the source URL, retrieval time, response status, raw response or permitted immutable snapshot, parser and selector versions, and transformation logs.
How often should a scraper run?
Match the schedule to the source’s change rate and your need for freshness, while staying within the operator’s limits.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




