Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Python Requests to download the HTTP response, then use an HTML/XML parser such as Beautiful Soup to extract data. A reliable scraper sets a descriptive User-Agent, uses a Session for related requests, applies an explicit connect/read timeout, checks status codes with raise_for_status(), and handles redirects, rate limits and transient failures. Requests cannot execute JavaScript, so content created only after a browser runs scripts requires an API or browser-capable capture tool instead.

What Requests does—and what it does not do

Requests is an HTTP client, not a scraper framework or HTML parser. It sends methods such as GET and POST, receives headers and a response body, and exposes that body through response.text, response.content or response.json(). Beautiful Soup parses HTML or XML and lets you select elements. The two libraries therefore solve different parts of the job.

The Requests project describes release 2.34.2 and official support for Python 3.10 and newer (documentation accessed in 2026). Beautiful Soup documentation reports version 4.14.3. Check the installed versions in your own environment before relying on a feature.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Requests: transport, headers, cookies, authentication, redirects, timeouts and connection reuse.
  • Beautiful Soup: HTML/XML parsing and selector-based extraction.
  • A browser or site API: JavaScript execution, rendered DOM state and interactions that do not exist in the initial HTTP response.

Install the libraries and make a first request

python -m pip install requests beautifulsoup4

Use a virtual environment for repeatable projects. The smallest useful request is:

import requests

url = "https://example.com/"
response = requests.get(
    url,
    headers={"User-Agent": "ExampleResearchBot/1.0 (+https://example.com/contact)"},
    timeout=(10, 30),
)
response.raise_for_status()
print(response.status_code)
print(response.url)          # final URL after permitted redirects
print(response.text[:500])

The two timeout values are a connection timeout and a read timeout. Requests applies no timeout unless you supply one. A read timeout limits waiting for data between bytes; it is not a guaranteed whole-download wall-clock deadline, so elapsed time can exceed the configured values.

Parse a page with Beautiful Soup

After the response succeeds, parse the body and validate selectors against representative pages. Prefer stable attributes over fragile positional selectors.

from bs4 import BeautifulSoup
import requests

url = "https://example.com/articles"
r = requests.get(url, timeout=(10, 30), headers={
    "User-Agent": "ExampleResearchBot/1.0 (+https://example.com/contact)"
})
r.raise_for_status()

soup = BeautifulSoup(r.text, "html.parser")
items = []
for article in soup.select("article.card"):
    title = article.select_one("h2")
    link = article.select_one("a[href]")
    if not title or not link:
        continue
    items.append({
        "title": title.get_text(" ", strip=True),
        "url": link.get("href"),
    })

for item in items:
    print(item)

Use html.parser for ordinary HTML. If the site returns XML, pass the appropriate parser available in your environment. Treat missing elements as normal: templates change, sponsored cards may differ, and an empty selector result should be logged rather than silently interpreted as a successful crawl.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a Session for cookies and repeated requests

A Session persists cookies and reuses connections to the same hosts. That matters for login flows, pagination and any crawl involving related requests.

import requests

with requests.Session() as session:
    session.headers.update({
        "User-Agent": "CatalogCollector/1.0 (+https://example.com/contact)",
        "Accept": "text/html,application/xhtml+xml",
    })
    session.get("https://example.com/login", timeout=(10, 30)).raise_for_status()
    page = session.get(
        "https://example.com/catalog",
        params={"page": 2, "sort": "newest"},
        timeout=(10, 30),
    )
    page.raise_for_status()
    print(page.url)
    print(page.text[:200])

Pass query data with params instead of manually concatenating and escaping a URL. For form submissions, use data=; for a JSON API, use json= and inspect response.json() only after checking that the response succeeded.

Check status codes and catch the right failures

Calling raise_for_status() turns 4xx and 5xx responses into an HTTPError. Catch the documented Requests exception family so a worker can record the URL, status (when available), retry count and failure class.

Symptom Likely meaning Practical response
ConnectionError DNS, refused connection, broken proxy or dropped network path. Check the host and proxy, then retry only when the failure is plausibly transient.
Timeout Connection or read phase exceeded its limit. Use separate connect/read values, log elapsed time and avoid unbounded retries.
HTTP 403 The server refused the request; it may require authentication or reject automated access. Do not attempt to evade controls. Verify permission, identify your client and use an official API where available.
HTTP 429 Rate limit reached. Honor Retry-After when supplied, reduce concurrency and increase spacing between requests.
HTTP 500–599 Server-side failure. Retry with bounded backoff for idempotent requests, then record the page as failed.
TooManyRedirects The redirect chain exceeded Requests’ limit. Inspect the URL and redirect policy; do not blindly follow a loop.
import requests

try:
    r = requests.get("https://example.com/data", timeout=(5, 20))
    r.raise_for_status()
except requests.exceptions.Timeout as exc:
    print("timeout", exc)
except requests.exceptions.TooManyRedirects as exc:
    print("redirect loop", exc)
except requests.exceptions.HTTPError as exc:
    status = exc.response.status_code if exc.response is not None else None
    print("http error", status, exc)
except requests.exceptions.ConnectionError as exc:
    print("connection failure", exc)
except requests.exceptions.RequestException as exc:
    print("other Requests failure", exc)

Retries, backoff, caching and rate limits

Retries are not a substitute for permission or capacity planning. Retry only operations that are safe to repeat, cap the number of attempts, and use increasing delays. A simple bounded loop is transparent:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import time
import requests

RETRYABLE = {429, 500, 502, 503, 504}

def get_with_backoff(session, url, attempts=4):
    for attempt in range(attempts):
        try:
            response = session.get(url, timeout=(10, 30))
            if response.status_code not in RETRYABLE:
                response.raise_for_status()
                return response
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after and retry_after.isdigit() else 2 ** attempt
        except (requests.exceptions.Timeout, requests.exceptions.ConnectionError):
            if attempt == attempts - 1:
                raise
            delay = 2 ** attempt
        if attempt == attempts - 1:
            response.raise_for_status()
        time.sleep(min(delay, 60))
    raise RuntimeError("unreachable")

Cache responses when freshness allows. A cache prevents duplicate downloads, lowers load on the target and makes reruns reproducible. Store the URL, retrieval time, status and body (or parsed result), and define an expiry policy rather than caching indefinitely.

Why a scraper hangs

The usual cause is an omitted timeout: a connection can wait indefinitely. Other causes include a slow origin, a proxy that accepts a connection but sends no bytes, a redirect chain, or downloading a very large response. Set both timeout components, log start and finish times, and stream unusually large files instead of holding them in memory.

with session.get(url, timeout=(5, 25), stream=True) as r:
    r.raise_for_status()
    with open("page.bin", "wb") as out:
        for chunk in r.iter_content(chunk_size=64 * 1024):
            if chunk:
                out.write(chunk)

A timeout does not cancel every activity at a single total deadline. If your job has a strict end-to-end budget, enforce that budget at the worker or job level as well as on each request.

Can Requests scrape JavaScript websites?

Only if the data you need is already in the initial HTTP response or exposed through a callable endpoint. Requests does not run the page’s JavaScript. Inspect the downloaded HTML: if the relevant records are absent and appear only after scripts execute, a Requests-only parser cannot produce them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an API when one exists

An official API usually provides a stable schema, authentication rules and a clearer usage limit than reverse-engineering browser calls. Prefer it for structured data and authenticated workflows.

Choose browser automation when rendering is essential

A browser-capable tool is appropriate when you must execute JavaScript, wait for a selector, click controls or capture the rendered state. It costs more resources and introduces browser, session and anti-bot considerations, so use it only for pages that need it.

Use Requests for the fast path

A practical crawler can first request the page and parse server-rendered content, then send only the pages that demonstrably require rendering to a browser or API. This keeps throughput and resource use predictable.

Responsible and lawful scraping

Technical success does not establish permission. Before crawling:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Read the site’s robots.txt and terms of service, and follow restrictions that apply to your use.
  • Identify your client honestly with a descriptive User-Agent and a contact URL where appropriate.
  • Limit concurrency and request frequency; treat 429 responses and Retry-After as instructions, not obstacles.
  • Collect only what you need, protect credentials and personal data, and set retention limits.
  • Use caching and conditional or incremental jobs where freshness permits.
  • Stop when access is denied or a bot check appears instead of trying to bypass it.

Rules differ by jurisdiction, site, data type and contract. Obtain legal advice for a commercial or high-volume project rather than treating a technical method as legal clearance.

A production-minded checklist

  1. Confirm that the target permits your intended collection and that an API is not the better interface.
  2. Install pinned, reviewed dependencies and record the Requests and Beautiful Soup versions.
  3. Create a Session, set a descriptive User-Agent and configure proxies or authentication deliberately.
  4. Set connect and read timeouts on every request.
  5. Call raise_for_status() before parsing and record the final URL and response status.
  6. Parse with explicit selectors and tests for missing or changed fields.
  7. Add bounded retries for transient failures, honoring Retry-After.
  8. Throttle, cache and checkpoint progress so a restart does not repeat the entire crawl.
  9. Log URL, timestamp, status, elapsed time, attempt number and failure class without logging secrets.
  10. Measure memory and response sizes; stream large bodies and cap unexpected downloads.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean screenshot or PDF rather than parsed records, ScreenshotNeo provides a single HTTP endpoint. It accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

For application code, the same call is:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also exposes an MCP server for AI clients such as Claude and Cursor, with take_screenshot, get_page_info and capture_pdf. Its 63 options include full-page and element capture, device presets, retina scale, dark mode, PDF paper and page ranges, custom CSS and JavaScript, clicks, selector waits, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Existing parameter names used by other screenshot APIs are accepted to ease switching.

Plan Included screenshots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free, and every feature is available on every plan. See the ScreenshotNeo documentation for parameters and response details. Start with 1,000 free screenshots a month—no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

How can I preserve the original bytes instead of decoded text?

Use response.content for the byte representation and write it in binary mode. Use response.text only when you want Requests’ decoded string.

How do I discover the encoding a server declared?

Inspect response.encoding before reading response.text. If the server declaration is wrong, set the correct encoding explicitly, then parse the text.

Should I run one process per URL?

Not necessarily. A Session can handle many related requests in one worker; choose concurrency from the site’s limits and your memory budget, then monitor failures rather than maximizing parallelism.

What should I keep for an auditable crawl?

Retain the input URL, retrieval timestamp, final URL, status, selected headers, parser version and extracted result, while excluding credentials and minimizing personal data.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.