Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Beautiful Soup parses HTML or XML markup into a Python object tree so you can search for tags, read attributes, extract text, and make changes. It is often used in web-scraping workflows, but it does not download web pages, run JavaScript, or crawl a site; your code must obtain the markup first.
What Beautiful Soup does
Beautiful Soup is a Python library for parsing HTML and XML. You give it markup—such as a string or an open file—and it builds a structured representation that Python can navigate. Instead of manually scanning a long text string, you can look for elements by tag, attribute, or other criteria, then retrieve their contents.
For example, this parses HTML that is already in a Python string:
from bs4 import BeautifulSoup
html = "<p class='notice'>Hello <b>Python</b></p>"
soup = BeautifulSoup(html, "html.parser")
notice = soup.find("p")
print(notice.get_text()) # Hello Python
The constructor turns the supplied markup into a tree. find() locates the first matching element, and get_text() returns its text. The official Beautiful Soup documentation describes the library as a way to pull data out of HTML and XML files.
#1 Best Overall
Where it fits in a scraping workflow
Web scraping usually involves distinct steps: request or otherwise obtain a page, parse its markup, extract the fields you need, and then store or use those fields. Beautiful Soup handles the parsing and navigation step. It does not make HTTP requests on its own. If you pass it a local file, it parses that file; if you want a live webpage, another component must fetch the response.
Here is a small end-to-end example using the Python requests package to fetch a page and Beautiful Soup to parse its HTML. Install the packages with python -m pip install requests beautifulsoup4.
import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
title = soup.find("h1")
print(title.get_text(" ", strip=True) if title else "No h1 found")
This example fetches the server’s HTTP response and parses the HTML in that response. It does not behave like a browser: pages that build or modify content with JavaScript may not include that content in the initial HTML. In that case, use an appropriate rendering approach to obtain the content before parsing, or use an API or data source offered by the site.
Common ways to find and extract data
Find one element or all matching elements
find() returns the first matching element, or None if there is no match. find_all() returns all matching elements. You can filter by tag name and attributes:
Rank #2
from bs4 import BeautifulSoup
html = """
<ul>
<li class="product" data-id="a1"><a href="/one">First</a></li>
<li class="product" data-id="b2"><a href="/two">Second</a></li>
</ul>
"""
soup = BeautifulSoup(html, "html.parser")
first_product = soup.find("li", class_="product")
products = soup.find_all("li", class_="product")
for product in products:
link = product.find("a")
print(product.get("data-id"), link.get("href"), link.get_text(strip=True))
Use class_ to search by the HTML class attribute; the underscore avoids colliding with Python’s reserved word. The get() method reads an attribute and returns None if it is absent.
Read text and attributes
For an element, get_text() gathers descendant text. Passing a separator can keep neighboring text pieces distinct, and strip=True removes leading and trailing whitespace. Attributes such as links’ href values are available through dictionary-style access or get().
link = soup.find("a")
if link:
label = link.get_text(" ", strip=True)
destination = link.get("href")
print(label, destination)
Check that an element exists before reading from it. Real pages change, and assuming a match exists can turn an ordinary missing field into an AttributeError.
Recommended Free Tools
Navigate the parsed tree
The parsed document is a tree of tags and text. You can move from a tag to its parent, children, or neighboring elements, or search within a specific part of the page. Keeping searches scoped to a relevant container can reduce accidental matches when a page repeats the same tag or class in navigation, recommendations, and main content.
Choosing a parser
Beautiful Soup provides a similar interface with several parser back ends, but they can build different trees when the input HTML is malformed. The parser choice affects speed, tolerance, dependencies, and reproducibility.
| Parser | When it fits | Trade-offs |
|---|---|---|
html.parser |
Basic use without adding a parser dependency; included with Python. | Reasonably fast, but less tolerant of malformed markup than html5lib and slower than lxml. |
lxml |
When speed matters. | Very fast according to the project’s qualitative guidance, but requires an external C dependency. |
html5lib |
When handling malformed HTML with browser-like parsing rules is important. | Highly tolerant, but slow and adds an external Python dependency. |
These are the project’s qualitative comparisons, not results from a new benchmark. Install optional parser packages only if you need them: python -m pip install lxml or python -m pip install html5lib. To get consistent behavior across machines, name the parser explicitly in the constructor rather than relying on whichever parser happens to be installed. For valid markup, parsers may give you the results you expect; malformed input is where their tree-building differences can matter.
Install and import the current package
For current Beautiful Soup 4 code, install the PyPI package named beautifulsoup4 and import it from bs4:
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11python -m pip install beautifulsoup4
from bs4 import BeautifulSoup
Do not install the similarly named PyPI package BeautifulSoup when following current Beautiful Soup 4 examples: the official project documentation identifies that package as the older Beautiful Soup 3 release. The current API documentation specifies Python 3.7 and later; Python 2 support ended on December 31, 2020, and Beautiful Soup 4.9.3 was the last release compatible with Python 2. See the project’s PyPI page for package information and its documentation for current usage.
What Beautiful Soup does not do
- It does not fetch a website. Provide markup from a request library, a file, or another source.
- It is not a browser or JavaScript runtime. Parsing HTML does not execute page scripts or render a dynamic page.
- It is not a crawler. Finding links in a document does not automatically visit them; your program would need to request pages and manage that workflow.
- It does not guarantee identical trees from every parser. For invalid markup, parser behavior can differ.
Troubleshooting common problems
ModuleNotFoundError: No module named 'bs4'
The package may not be installed in the Python environment running your script. Install it using that interpreter: python -m pip install beautifulsoup4. If you use a virtual environment or a different Python command, activate the environment or substitute its interpreter when running pip.
Beautiful Soup says a parser is unavailable
If you selected lxml or html5lib, install that optional dependency in the same environment as your script. Alternatively, use Python’s built-in html.parser where its behavior is sufficient.
find() returns None
The tag or attribute may not appear in the markup you supplied, the page may have changed, or the content may be added by JavaScript after the initial response. Inspect the actual response or file first, then adjust the selector or obtain rendered content by another method. Check for None before accessing a result’s attributes or methods.
Extracted text is missing or oddly spaced
Inspect the parsed element and its descendants: the text may be in another part of the document, separated across nested tags, or absent from the source HTML. Try get_text(" ", strip=True) to make text boundaries and whitespace easier to handle. If the expected text only appears after scripts run in a browser, parsing the original response will not create it.
Best Value
Results change between computers
Specify a parser explicitly and ensure the same parser dependency is installed in each environment. Different parser implementations can interpret malformed HTML differently, so compare the input markup and parser configuration when results diverge.
Or skip the browser setup
If what you need is a clean screenshot or PDF rather than a parsed data tree, ScreenshotNeo is a website screenshot API and MCP server. A single request can return PNG, JPEG, WebP, or PDF; its options include full-page capture, selected elements, device and viewport settings, custom CSS or JavaScript, and waiting for a selector or network idle. Its capture workflow accepts cookie banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, with response headers indicating the page verdict and billing status. AI agents can use its MCP server tools for screenshots, page info, and PDF capture. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Which tool should you use?
Use Beautiful Soup when you need to inspect markup and extract structured data in Python. Pair it with a separate HTTP client when the source is a web response. Choose a browser-based capture or rendering tool when the output you need is a visual screenshot, PDF, or content that depends on browser execution. These tools solve different steps: a parsed HTML tree is not a screenshot, and a screenshot does not give you Beautiful Soup’s navigable tag structure.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

