October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk7 min

How to Use Beautiful Soup for Web Scraping in Python

Beautiful Soup parses HTML; Requests fetches it. Follow a practical Python 3 guide to parser choice, element searches, extraction, and troubleshooting.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Beautiful Soup parses HTML or XML that you give it; it does not download web pages or run their JavaScript. A basic scraper therefore has two jobs: retrieve the page with an HTTP client such as Requests, then parse the returned markup with Beautiful Soup. This guide shows the complete Python 3 workflow, how to choose a parser, find and extract data, and diagnose common failures.

What Beautiful Soup does—and what it does not do

Beautiful Soup turns supplied HTML or XML into a navigable tree of Python objects. It can search that tree, extract text and attributes, and modify the parsed structure. The project describes it as “a Python library for pulling data out of HTML and XML files.” Beautiful Soup documentation

As an Amazon Associate I earn from qualifying purchases.

It does not send HTTP requests, crawl a site, or execute browser JavaScript. Use an HTTP client to retrieve markup, check what came back, and then pass the response content to Beautiful Soup. If a page is built in the browser after scripts run, the initial HTTP response may not contain the data you see on screen.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Beautiful Soup and Requests

Use Python 3. Install the Beautiful Soup package and Requests, which the example uses to fetch a page:

python -m pip install beautifulsoup4 requests

The package is named beautifulsoup4, but you import it from the bs4 namespace. The older Python 2-specific installation path is not relevant to this Python 3 example. Beautiful Soup on PyPI

Fetch a page, parse it, and extract data

Save this as scrape.py and run it with Python. It checks the HTTP response before parsing, explicitly selects a parser, handles a missing title, and extracts links that actually have an href attribute.

import requests
from bs4 import BeautifulSoup

url = "https://example.com/"

response = requests.get(
    url,
    headers={"User-Agent": "Example scraper contact: [email protected]"},
    timeout=20,
)
response.raise_for_status()

soup = BeautifulSoup(response.content, "html.parser")

title = soup.title.get_text(strip=True) if soup.title else "(no title found)"
print("Title:", title)

for link in soup.find_all("a", href=True):
    text = link.get_text(" ", strip=True)
    href = link.get("href")
    print(text, href)

Replace the example URL and contact text with values appropriate for your use. Requests’ quickstart documents its response object and the process of making requests. Requests Quickstart

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why pass response.content?

response.content is the response body as bytes. Beautiful Soup can use that input while accounting for encoding information in the markup. Requests also provides response.text, a decoded string whose encoding is selected from the response information and detection. If characters look corrupted, inspect the response headers and encoding before changing your selectors.

Check what the server returned

raise_for_status() stops the script when the response has an unsuccessful HTTP status instead of silently parsing an error page as if it were the intended document. For unexpected results, inspect response.status_code, response.headers, and a portion of response.text or response.content before debugging the parse.

Choose a parser deliberately

Beautiful Soup supports Python’s built-in html.parser and optional parsers including lxml and html5lib. The parser interprets markup and constructs the tree Beautiful Soup searches. Malformed HTML can produce different trees with different parsers, so specify the parser instead of relying on an environment-dependent default.

Parser When to choose it Consideration
html.parser A built-in option for HTML with no extra parser package. Explicitly name it in the constructor for consistent intent.
lxml When you have installed and chosen the lxml parser; use its XML mode for XML. It is an optional dependency. Parser choice can affect the resulting tree.
html5lib When you have installed and chosen html5lib for parsing HTML. It is optional, and can produce a different tree from other parsers.

Install an optional parser separately if you choose it, then name it in the constructor, for example BeautifulSoup(response.content, "lxml"). For XML, use BeautifulSoup(xml_bytes, "xml") with lxml installed, as directed in the Beautiful Soup documentation. The documentation options establish different parser choices; they do not establish current speed rankings, so choose for input compatibility and predictable output rather than an assumed benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find elements and extract their values

Use find() for one match

find() returns the first matching element or None. Check for a match before accessing its text or attributes:

heading = soup.find("h1")
if heading is not None:
    print(heading.get_text(" ", strip=True))

Use find_all() for repeated matches

find_all() returns all matching elements. You can filter by tag name and attributes, then extract text or an attribute from each tag:

for item in soup.find_all("a", class_="product-link"):
    label = item.get_text(" ", strip=True)
    destination = item.get("href")
    print(label, destination)

In Python, class_ is used because class is a reserved word. The attribute lookup with get() safely returns None if the attribute is absent.

Use CSS selectors when they clarify relationships

select() accepts CSS selectors and is useful when a relationship or combination of attributes is clearer as a selector:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
for item in soup.select("article.product a.product-link[href]"):
    print(item.get_text(" ", strip=True), item.get("href"))

Use whichever form makes the target structure easiest to understand and maintain. Verify the selector against the markup you actually fetched; do not assume a page’s visual arrangement or an element’s position will remain stable.

Extract text without assuming a match exists

For a tag, get_text(" ", strip=True) combines descendant text, separates pieces with spaces, and strips surrounding whitespace. A missing element is not an empty tag, so test the lookup result before calling methods on it. Avoid brittle assumptions such as treating the third paragraph as a price unless the target page explicitly guarantees that structure.

When the returned HTML does not match the browser

A simple Requests-and-Beautiful-Soup script parses the HTTP response; it does not render the page as a browser or execute page scripts. A site may insert content after JavaScript runs, so the browser’s visible text can be absent from the response. Inspect the response body first. If the required content is not present there, changing Beautiful Soup selectors will not make it appear: use a retrieval method appropriate to the page, subject to the site’s rules.

Troubleshoot common scraping failures

Symptom Likely cause What to check or do
Lookup returns None or an empty list The selector does not match the returned markup, the page structure changed, or the data is added by JavaScript. Inspect the response body and parsed tree; verify tag names, attributes, and nesting. If the content is not in the response, the simple request-and-parse approach cannot extract it.
The script parses an error page or unexpected document The HTTP request returned an unsuccessful response or different content than expected. Check the status code, headers, and response body; call raise_for_status() before parsing.
Text contains broken or unexpected characters The response encoding or decoding may not match the content. Inspect response headers and encoding, and compare response.text with the raw response.content bytes before altering selectors.
Results differ across computers A different parser may have constructed a different tree from imperfect markup. Specify the parser explicitly and ensure the chosen optional parser is installed in each environment.
A tag exists but the expected attribute is missing The matching element may not have that attribute. Use tag.get("attribute") and handle a None result instead of assuming the attribute exists.

Keep scraping responsible and maintainable

  • Check the target site’s current terms, access controls, robots directives, and other applicable requirements before collecting data; obligations can differ by site and jurisdiction.
  • Obtain authorization where needed and avoid sending requests at a rate that burdens the service.
  • Keep selectors tied to meaningful tags and attributes, and revisit them when the site’s markup changes.
  • Log enough response and extraction information to distinguish an HTTP problem from a parsing mismatch.

Beautiful Soup and Requests documentation describe software behavior, not whether scraping a particular site is permitted. Make that determination for the specific site and your circumstances.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your task is to capture a rendered page rather than parse HTML yourself, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF; its clean-shot options accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Each step can be turned off.

For example, this saves a WebP screenshot of the target URL. See the ScreenshotNeo API documentation for request options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo bills only clean shots: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents, including Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, with no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does Beautiful Soup scrape a page on its own?

No. An HTTP client retrieves the page; Beautiful Soup parses the markup it receives.

Can Beautiful Soup extract content that appears only after JavaScript runs?

Not from a response that does not contain that content. Beautiful Soup parses markup and does not execute page scripts.

Should I use find_all() or select()?

Use the method that expresses the target structure most clearly and is easiest for you to maintain; verify it against the fetched markup.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.