Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Python Requests to download the HTTP response, then use an HTML/XML parser such as Beautiful Soup to extract data. A reliable scraper sets a descriptive User-Agent, uses a Session for related requests, applies an explicit connect/read timeout, checks status codes with raise_for_status(), and handles redirects, rate limits and transient failures. Requests cannot execute JavaScript, so content created only after a browser runs scripts requires an API or browser-capable capture tool instead.
What Requests does—and what it does not do
Requests is an HTTP client, not a scraper framework or HTML parser. It sends methods such as GET and POST, receives headers and a response body, and exposes that body through response.text, response.content or response.json(). Beautiful Soup parses HTML or XML and lets you select elements. The two libraries therefore solve different parts of the job.
The Requests project describes release 2.34.2 and official support for Python 3.10 and newer (documentation accessed in 2026). Beautiful Soup documentation reports version 4.14.3. Check the installed versions in your own environment before relying on a feature.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →- Requests: transport, headers, cookies, authentication, redirects, timeouts and connection reuse.
- Beautiful Soup: HTML/XML parsing and selector-based extraction.
- A browser or site API: JavaScript execution, rendered DOM state and interactions that do not exist in the initial HTTP response.
Install the libraries and make a first request
python -m pip install requests beautifulsoup4
Use a virtual environment for repeatable projects. The smallest useful request is:
#1 Best Overall
import requests
url = "https://example.com/"
response = requests.get(
url,
headers={"User-Agent": "ExampleResearchBot/1.0 (+https://example.com/contact)"},
timeout=(10, 30),
)
response.raise_for_status()
print(response.status_code)
print(response.url) # final URL after permitted redirects
print(response.text[:500])
The two timeout values are a connection timeout and a read timeout. Requests applies no timeout unless you supply one. A read timeout limits waiting for data between bytes; it is not a guaranteed whole-download wall-clock deadline, so elapsed time can exceed the configured values.
Parse a page with Beautiful Soup
After the response succeeds, parse the body and validate selectors against representative pages. Prefer stable attributes over fragile positional selectors.
from bs4 import BeautifulSoup
import requests
url = "https://example.com/articles"
r = requests.get(url, timeout=(10, 30), headers={
"User-Agent": "ExampleResearchBot/1.0 (+https://example.com/contact)"
})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
items = []
for article in soup.select("article.card"):
title = article.select_one("h2")
link = article.select_one("a[href]")
if not title or not link:
continue
items.append({
"title": title.get_text(" ", strip=True),
"url": link.get("href"),
})
for item in items:
print(item)
Use html.parser for ordinary HTML. If the site returns XML, pass the appropriate parser available in your environment. Treat missing elements as normal: templates change, sponsored cards may differ, and an empty selector result should be logged rather than silently interpreted as a successful crawl.
Free tools Windows power users keep installed
One-click scans. No signup required.
Use a Session for cookies and repeated requests
A Session persists cookies and reuses connections to the same hosts. That matters for login flows, pagination and any crawl involving related requests.
import requests
with requests.Session() as session:
session.headers.update({
"User-Agent": "CatalogCollector/1.0 (+https://example.com/contact)",
"Accept": "text/html,application/xhtml+xml",
})
session.get("https://example.com/login", timeout=(10, 30)).raise_for_status()
page = session.get(
"https://example.com/catalog",
params={"page": 2, "sort": "newest"},
timeout=(10, 30),
)
page.raise_for_status()
print(page.url)
print(page.text[:200])
Pass query data with params instead of manually concatenating and escaping a URL. For form submissions, use data=; for a JSON API, use json= and inspect response.json() only after checking that the response succeeded.
Rank #2
Check status codes and catch the right failures
Calling raise_for_status() turns 4xx and 5xx responses into an HTTPError. Catch the documented Requests exception family so a worker can record the URL, status (when available), retry count and failure class.
| Symptom | Likely meaning | Practical response |
|---|---|---|
ConnectionError |
DNS, refused connection, broken proxy or dropped network path. | Check the host and proxy, then retry only when the failure is plausibly transient. |
Timeout |
Connection or read phase exceeded its limit. | Use separate connect/read values, log elapsed time and avoid unbounded retries. |
| HTTP 403 | The server refused the request; it may require authentication or reject automated access. | Do not attempt to evade controls. Verify permission, identify your client and use an official API where available. |
| HTTP 429 | Rate limit reached. | Honor Retry-After when supplied, reduce concurrency and increase spacing between requests. |
| HTTP 500–599 | Server-side failure. | Retry with bounded backoff for idempotent requests, then record the page as failed. |
TooManyRedirects |
The redirect chain exceeded Requests’ limit. | Inspect the URL and redirect policy; do not blindly follow a loop. |
import requests
try:
r = requests.get("https://example.com/data", timeout=(5, 20))
r.raise_for_status()
except requests.exceptions.Timeout as exc:
print("timeout", exc)
except requests.exceptions.TooManyRedirects as exc:
print("redirect loop", exc)
except requests.exceptions.HTTPError as exc:
status = exc.response.status_code if exc.response is not None else None
print("http error", status, exc)
except requests.exceptions.ConnectionError as exc:
print("connection failure", exc)
except requests.exceptions.RequestException as exc:
print("other Requests failure", exc)
Retries, backoff, caching and rate limits
Retries are not a substitute for permission or capacity planning. Retry only operations that are safe to repeat, cap the number of attempts, and use increasing delays. A simple bounded loop is transparent:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsimport time
import requests
RETRYABLE = {429, 500, 502, 503, 504}
def get_with_backoff(session, url, attempts=4):
for attempt in range(attempts):
try:
response = session.get(url, timeout=(10, 30))
if response.status_code not in RETRYABLE:
response.raise_for_status()
return response
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after and retry_after.isdigit() else 2 ** attempt
except (requests.exceptions.Timeout, requests.exceptions.ConnectionError):
if attempt == attempts - 1:
raise
delay = 2 ** attempt
if attempt == attempts - 1:
response.raise_for_status()
time.sleep(min(delay, 60))
raise RuntimeError("unreachable")
Cache responses when freshness allows. A cache prevents duplicate downloads, lowers load on the target and makes reruns reproducible. Store the URL, retrieval time, status and body (or parsed result), and define an expiry policy rather than caching indefinitely.
Why a scraper hangs
The usual cause is an omitted timeout: a connection can wait indefinitely. Other causes include a slow origin, a proxy that accepts a connection but sends no bytes, a redirect chain, or downloading a very large response. Set both timeout components, log start and finish times, and stream unusually large files instead of holding them in memory.
with session.get(url, timeout=(5, 25), stream=True) as r:
r.raise_for_status()
with open("page.bin", "wb") as out:
for chunk in r.iter_content(chunk_size=64 * 1024):
if chunk:
out.write(chunk)
A timeout does not cancel every activity at a single total deadline. If your job has a strict end-to-end budget, enforce that budget at the worker or job level as well as on each request.
Can Requests scrape JavaScript websites?
Only if the data you need is already in the initial HTTP response or exposed through a callable endpoint. Requests does not run the page’s JavaScript. Inspect the downloaded HTML: if the relevant records are absent and appear only after scripts execute, a Requests-only parser cannot produce them.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Choose an API when one exists
An official API usually provides a stable schema, authentication rules and a clearer usage limit than reverse-engineering browser calls. Prefer it for structured data and authenticated workflows.
Choose browser automation when rendering is essential
A browser-capable tool is appropriate when you must execute JavaScript, wait for a selector, click controls or capture the rendered state. It costs more resources and introduces browser, session and anti-bot considerations, so use it only for pages that need it.
Use Requests for the fast path
A practical crawler can first request the page and parse server-rendered content, then send only the pages that demonstrably require rendering to a browser or API. This keeps throughput and resource use predictable.
Responsible and lawful scraping
Technical success does not establish permission. Before crawling:
- Read the site’s
robots.txtand terms of service, and follow restrictions that apply to your use. - Identify your client honestly with a descriptive User-Agent and a contact URL where appropriate.
- Limit concurrency and request frequency; treat 429 responses and
Retry-Afteras instructions, not obstacles. - Collect only what you need, protect credentials and personal data, and set retention limits.
- Use caching and conditional or incremental jobs where freshness permits.
- Stop when access is denied or a bot check appears instead of trying to bypass it.
Rules differ by jurisdiction, site, data type and contract. Obtain legal advice for a commercial or high-volume project rather than treating a technical method as legal clearance.
A production-minded checklist
- Confirm that the target permits your intended collection and that an API is not the better interface.
- Install pinned, reviewed dependencies and record the Requests and Beautiful Soup versions.
- Create a Session, set a descriptive User-Agent and configure proxies or authentication deliberately.
- Set connect and read timeouts on every request.
- Call
raise_for_status()before parsing and record the final URL and response status. - Parse with explicit selectors and tests for missing or changed fields.
- Add bounded retries for transient failures, honoring
Retry-After. - Throttle, cache and checkpoint progress so a restart does not repeat the entire crawl.
- Log URL, timestamp, status, elapsed time, attempt number and failure class without logging secrets.
- Measure memory and response sizes; stream large bodies and cap unexpected downloads.
Or skip the browser setup
If your goal is a clean screenshot or PDF rather than parsed records, ScreenshotNeo provides a single HTTP endpoint. It accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
For application code, the same call is:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also exposes an MCP server for AI clients such as Claude and Cursor, with take_screenshot, get_page_info and capture_pdf. Its 63 options include full-page and element capture, device presets, retina scale, dark mode, PDF paper and page ranges, custom CSS and JavaScript, clicks, selector waits, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Existing parameter names used by other screenshot APIs are accepted to ease switching.
Best Value
| Plan | Included screenshots | Price |
|---|---|---|
| Free | 1,000 per month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing gives two months free, and every feature is available on every plan. See the ScreenshotNeo documentation for parameters and response details. Start with 1,000 free screenshots a month—no card required.
Frequently Asked Questions
How can I preserve the original bytes instead of decoded text?
Use response.content for the byte representation and write it in binary mode. Use response.text only when you want Requests’ decoded string.
How do I discover the encoding a server declared?
Inspect response.encoding before reading response.text. If the server declaration is wrong, set the correct encoding explicitly, then parse the text.
Should I run one process per URL?
Not necessarily. A Session can handle many related requests in one worker; choose concurrency from the site’s limits and your memory budget, then monitor failures rather than maximizing parallelism.
What should I keep for an auditable crawl?
Retain the input URL, retrieval timestamp, final URL, status, selected headers, parser version and extracted result, while excluding credentials and minimizing personal data.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

