Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Requests to download a page, check that the HTTP response is usable, then give its HTML to Beautiful Soup for searching and extraction. This two-library workflow is reliable when the data you need is present in the server-returned markup. It is not a promise that every site can be collected with plain HTTP: pages that render data in JavaScript, require authentication, block automated clients, or prohibit collection need a different approach and careful review of their rules.

How do I use Beautiful Soup with Requests?

Requests and Beautiful Soup have separate jobs. Requests is the HTTP client: it sends a GET request, receives headers and a body, and exposes status, text, bytes, and other response data. Beautiful Soup is the parser: it turns that markup into a navigable tree whose tags, attributes, text, and relationships you can search. The Requests project describes its library as “An elegant and simple HTTP library for Python, built for human beings” (official Requests documentation).

Install the packages in the Python environment that will run your script. The current Requests overview states support for Python 3.10 and newer; verify the requirement again when you publish or deploy because support can change.

python -m pip install requests beautifulsoup4

The package is named beautifulsoup4, but the import is from bs4 import BeautifulSoup. This first example uses Python’s built-in html.parser, so it does not add another parser dependency:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
from bs4 import BeautifulSoup

url = "https://example.com/"
response = requests.get(url, timeout=(10, 30))
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
print(soup.title.get_text(strip=True) if soup.title else "No title")

for link in soup.select("a[href]"):
    print(link.get_text(" ", strip=True), link["href"])

The connect/read timeout tuple limits how long connection setup and response reading may wait. A single float, such as timeout=30, is also accepted. Calling raise_for_status() before parsing prevents a 404, 500, or similar response from being mistaken for the page you wanted. A 200 status still does not prove that the expected content is present, so inspect the extracted result as well.

Keep TLS certificate verification enabled, which Requests does by default. The API documentation warns that verify=False accepts unverified certificates and can expose an application to man-in-the-middle attacks; do not use it in ordinary production code.

How do I scrape a webpage with Python?

1. Define the fields and check the returned document

Write down the values you actually need—such as a heading, product name, price, or article links. Fetch one page first and inspect response.status_code, response.url, and a small portion of the body. A successful request can return a login page, an error template, or HTML without the data you expected.

print(response.status_code, response.url)
print(response.headers.get("content-type"))
print(response.text[:500])

2. Parse with an explicitly selected backend

Beautiful Soup accepts several parser implementations. The explicit second argument makes your choice visible and helps explain differences between machines:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

soup = BeautifulSoup(response.content, "html.parser")

Use response.content (original bytes) when you need to investigate encoding. Requests guesses an encoding from HTTP headers and available detection libraries; you can inspect or set response.encoding before reading response.text. Beautiful Soup converts parsed markup to Unicode for normal text work.

3. Locate elements by tag, attributes, and CSS selector

Beautiful Soup supports direct methods, tree navigation, and CSS selectors through its SoupSieve integration. Selector support follows the installed Beautiful Soup/SoupSieve versions, so consult the documentation for the versions in your environment.

# First matching element
headline = soup.find("h1")

# All matching elements with an attribute
cards = soup.find_all("article", class_="card")

# CSS selectors
for item in soup.select("ul.results > li[data-id]"):
    identifier = item.get("data-id")
    name = item.select_one(".name")
    print(identifier, name.get_text(" ", strip=True) if name else None)

4. Extract text and attributes defensively

Use get_text(" ", strip=True) to collapse nested text into readable output. Use tag.get("attribute") when an attribute may be absent, rather than indexing it unconditionally.

records = []
for card in soup.select("article.card"):
    title = card.select_one("h2, h3")
    price = card.select_one(".price")
    image = card.select_one("img[src]")
    records.append({
        "title": title.get_text(" ", strip=True) if title else None,
        "price": price.get_text(" ", strip=True) if price else None,
        "image_url": image.get("src") if image else None,
    })

5. Verify against real markup and save structured output

Selectors describe the HTML you received, not a permanent API. Check a sample manually, handle missing fields, and log the source URL. Serialize only after validation:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json

with open("records.json", "w", encoding="utf-8") as file:
    json.dump(records, file, ensure_ascii=False, indent=2)

Which parser should I use with Beautiful Soup?

Parser Typical characteristics in the Beautiful Soup guide Choose it when
html.parser Built in with Python; described as reasonably fast and requiring no extra package. You want a simple example or minimal dependency footprint.
lxml Described as very fast and lenient, with an external C dependency. Your deployment permits the dependency and throughput matters.
html5lib Described as very lenient and browser-like, but slower and requiring an external Python package. HTML5-style recovery of badly malformed documents is more important than speed.

Install and select the backend explicitly, for example:

python -m pip install lxml html5lib

soup_lxml = BeautifulSoup(response.content, "lxml")
soup_html5 = BeautifulSoup(response.content, "html5lib")

Invalid HTML can produce different trees with different parsers. That changes which element a selector finds, so name the backend in your code and pin compatible dependencies when repeatable output matters. The descriptions above are documented trade-offs, not universal benchmark results; measure a representative workload before choosing on speed alone.

Why is Beautiful Soup not finding my element?

The element is created by JavaScript

Requests receives the server response; it does not execute browser JavaScript. If the data appears only after a script runs, it will not be in response.text. Inspect the returned HTML and, where permitted, identify an official data endpoint or use a browser automation workflow suited to the site.

The selector does not match the returned markup

Class names, nesting, and attributes may differ from what you saw in developer tools. Save a response sample, search it for a distinctive string, and test a narrow selector with soup.select_one(). Avoid assuming a browser’s live DOM is identical to the original response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The request returned a different page

Check status, final URL, content type, redirects, and whether a login, consent, bot-check, or rate-limit page was returned. Add appropriate headers only when you have a legitimate reason and permission; do not attempt to bypass access controls.

Encoding is wrong

Look at response.headers.get("content-type") and response.apparent_encoding where available. If the declared encoding is incorrect, set it before accessing text, or parse the original bytes:

response.encoding = "utf-8"
soup = BeautifulSoup(response.text, "html.parser")

The parser changed the tree

Malformed markup is repaired differently by html.parser, lxml, and html5lib. Compare the same input with your explicitly selected backend and install that backend everywhere the scraper runs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reliability, performance, and responsible collection

  • Use timeouts, check status, and catch targeted exceptions such as requests.exceptions.Timeout and requests.exceptions.RequestException.
  • Reuse a requests.Session() for multiple pages when appropriate; it can preserve connections and shared headers.
  • Throttle requests, honor stated rate limits, and cache responses during development.
  • Keep selectors and parser versions under test. A small fixture HTML file can detect template changes before a scheduled job produces bad data.
  • Use bytes when diagnosing encoding; use text after confirming or correcting the encoding.
  • Review the target site’s terms, robots guidance, authentication requirements, data rights, and applicable legal requirements. Library documentation explains mechanics, not permission for a particular target or jurisdiction.

For transient failures, retry only safe requests with bounded backoff and a maximum attempt count. Do not blindly retry authentication failures, client errors, or a site that is explicitly refusing automated traffic.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than extracting fields from HTML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF; it accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status.

For a direct capture, see the ScreenshotNeo documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python and Node.js clients can use the same endpoint:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It also offers full-page and element captures, device and retina settings, PDFs, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and an MCP server with take_screenshot, get_page_info, and capture_pdf for AI clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Further reading

A Python web scraping book can provide longer exercises, but it is optional; Requests and Beautiful Soup are free libraries. Verify any specific title, edition, price, and availability before buying.

Frequently Asked Questions

Can Requests and Beautiful Soup scrape every website?

No. They work best when the needed content is in the HTTP response HTML. JavaScript-rendered content, authentication, bot checks, rate limits, and site rules may require another workflow or make collection inappropriate.

Should I parse response.text or response.content?

Use response.text after checking or correcting the detected encoding. Use response.content when you need the original bytes or are diagnosing an encoding problem.

Is a 200 response proof that scraping succeeded?

No. It only indicates the server returned a successful HTTP status. Confirm the page content and required fields are actually present.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.