October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk7 min

How to Scrape Websites with Beautiful Soup in Python

Beautiful Soup parses HTML but does not fetch it. Pair it with Requests to retrieve a page, check the response, and extract data safely.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Beautiful Soup parses HTML; it does not download web pages. To scrape a page, use an HTTP client such as Requests to fetch its response, check that the request succeeded, then parse the returned markup with Beautiful Soup and extract the fields you need. The example below retrieves links, handles missing attributes, and includes the checks that prevent common empty-result and hanging-request problems.

What Beautiful Soup does—and what it does not

Beautiful Soup turns HTML or XML into a navigable tree. You can search that tree by tag, attribute, or CSS selector and read text or attributes from matching elements. It is a parser, not a web client: pair it with Requests for a remote page, or pass it HTML you already have.

As an Amazon Associate I earn from qualifying purchases.

A scraper can only inspect the markup its HTTP client receives. When a page fills in content later with JavaScript, that content may not appear in the initial response, so parsing the response alone may not find it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the packages

Install Beautiful Soup 4 using the distribution name beautifulsoup4; import it in Python as bs4. This example also uses Requests and Python’s built-in html.parser.

python -m pip install beautifulsoup4 requests

Run the command in the same Python environment that will run your script. If you prefer another parser, install its package explicitly and use its parser name in BeautifulSoup.

Fetch a page, check it, and extract links

This complete example fetches a page with a timeout, raises an error for unsuccessful HTTP responses, parses the returned HTML, and collects anchors that actually have an href attribute.

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

url = "https://example.com/"

try:
    response = requests.get(url, timeout=20)
    response.raise_for_status()
except requests.exceptions.Timeout:
    raise SystemExit(f"The request to {url} timed out")
except requests.exceptions.RequestException as exc:
    raise SystemExit(f"Could not fetch {url}: {exc}")

soup = BeautifulSoup(response.text, "html.parser")

links = []
for anchor in soup.find_all("a", href=True):
    href = anchor.get("href")
    text = anchor.get_text(" ", strip=True)
    links.append({
        "text": text,
        "url": urljoin(response.url, href),
    })

print(f"HTTP status: {response.status_code}")
print(f"Links found: {len(links)}")
for link in links:
    print(link["text"], link["url"])

Replace https://example.com/ with a page you are allowed to access. response.raise_for_status() prevents an error response from being silently treated like a successful page. Requests does not set a timeout unless you provide one; its Quickstart recommends using the timeout parameter in nearly all production requests. The example uses a single timeout value; Requests also supports separate connect and read timeout values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

response.text is the decoded response body. If you need to inspect what the server actually returned, print a short slice of it or save it locally before deciding that the selector is wrong. urljoin converts relative link paths into absolute URLs using the final response URL.

Choose the right Beautiful Soup search method

Use find() for one match

find() returns the first matching tag, or None if no match exists. Check the result before accessing its attributes:

heading = soup.find("h1")
if heading is not None:
    print(heading.get_text(" ", strip=True))
else:
    print("No h1 found")

Use find_all() for repeated elements

find_all() returns all matches; no match is a valid empty result. Attribute filters are useful when the markup has a stable identifier:

for item in soup.find_all("div", class_="product-card"):
    title = item.find("h2")
    if title is not None:
        print(title.get_text(" ", strip=True))

Filters can target tag names and attributes, and Beautiful Soup supports other filter forms such as regular expressions, lists, functions, and True. For a potentially absent attribute, use tag.get("href") or filter for its presence with href=True.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use CSS selectors when they express the target clearly

select() returns all matches for a CSS selector, while select_one() returns the first match or None. For example:

cards = soup.select("article.product-card")
first_title = soup.select_one("article.product-card h2")

Selector support depends on the installed Beautiful Soup and selector-engine versions. The Beautiful Soup documentation describes support for most CSS4 selectors through SoupSieve in modern versions; check the documentation and your installed environment if a selector does not behave as expected.

Inspect the markup before writing selectors

  1. Fetch and validate. Check the HTTP status and make sure the response is the expected page, not an error or challenge page.
  2. Inspect the response HTML. Search its source for a visible phrase or element you expected to find. The browser’s rendered view can differ from the raw HTTP response.
  3. Identify stable structure. Prefer meaningful tags and attributes, such as an article class or link destination, over brittle positional assumptions.
  4. Try a small extraction. Print a few matches and their attributes before building a larger scraper around them.
  5. Handle absence deliberately. Check for None after find(), and decide what an empty find_all() result should mean for your task.

Choose a parser for repeatable results

Beautiful Soup can build its tree with html.parser, lxml, or html5lib. Different parsers can produce different trees from malformed HTML, so specify the parser rather than relying on an implicit choice when repeatability matters.

Parser Practical note
html.parser Built into Python; no separate parser package is required.
lxml A third-party option that the Beautiful Soup documentation describes as faster. Install it in the active environment before passing "lxml".
html5lib A third-party option the documentation describes as parsing more like a browser. Install it in the active environment before passing "html5lib".

The cited Beautiful Soup documentation page says it covers version 4.8.1, so verify current package and selector details against the version installed in your environment. If raw parsing speed is the overriding concern, the Beautiful Soup documentation recommends using lxml directly rather than its convenience layer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If you need a rendered screenshot or PDF rather than structured text extracted from HTML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF; it is not a replacement for Beautiful Soup when you need to parse page data.

For example, save a screenshot of a page as WebP with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options and response details. Cookie banners and consent overlays, newsletter popups, and chat widgets can be removed before the capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. The MCP server includes take_screenshot, get_page_info, and capture_pdf for AI agents. The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common problems

Beautiful Soup returns an empty list

  • Confirm the fetch succeeded and the response body contains the expected content; check the status and inspect a snippet of the HTML.
  • Re-check the selector against the current response markup. A site may have changed its tags, classes, or attributes.
  • If the content appears in a browser but not in the response, it may be populated after the initial HTML arrives. Beautiful Soup cannot parse content that is absent from the input it receives.
  • Remember that find_all() returns an empty result when nothing matches; that result does not itself indicate a parser error.

You get a NoneType or attribute error

find() can return None. Test for a match before calling .get(), reading text, or accessing another tag property.

The request hangs or you parse an error page

Provide a timeout and call raise_for_status(). Handle Requests exceptions so timeouts and other request failures are visible rather than mistaken for extraction results. Inspect status and response content when a server returns an unexpected page.

The same HTML produces different parse results

Choose and name a parser explicitly, confirm its dependency is installed in the active environment, and keep the parser choice consistent. Different parsers may repair malformed markup differently.

ImportError or the wrong package name

Install beautifulsoup4 and import from bs4. If installation appears successful but the import fails, check that python -m pip refers to the same Python environment running your script.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep scraping maintainable and considerate

  • Review the specific site’s current terms and access guidance before collecting data; this general guide does not establish permission for any particular site.
  • Keep request volume modest and avoid collecting personal data you do not need.
  • Stop if the site blocks access instead of trying to evade the block.
  • Keep extraction assumptions visible in code and validate results, since page structure can change.

Frequently Asked Questions

Can Beautiful Soup scrape a website without Requests?

Yes, if you already have the HTML—for example, from a file or another HTTP client. Beautiful Soup parses input; it does not fetch remote pages itself.

Does Beautiful Soup run JavaScript?

No. It parses the markup it is given. If the content you need is added after the initial response, it may not be present in that markup.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.