To parse web data with Python and Beautiful Soup, first obtain the page’s HTML, then pass that markup to BeautifulSoup with an explicit parser. Find the elements you need with find(), find_all(), or CSS selectors, and extract text or attributes from the matches. Beautiful Soup parses markup; it does not fetch a page or run its JavaScript.
What Beautiful Soup parses—and what it does not
Beautiful Soup turns HTML or XML markup into a Python object tree that you can navigate, search, and modify. It is a parsing library, not an HTTP client or browser. Keep the workflow in two parts: acquire the response HTML from a source you are allowed to access, then parse the response body.
This distinction matters when a page looks different in a browser than it does in your script. The HTML received over HTTP may not contain content that a site adds later with JavaScript. Beautiful Soup can only inspect the markup supplied to it; it cannot retrieve content that is absent from that markup. The official documentation describes its parsing and search APIs at Beautiful Soup Documentation, while Python documents URL-opening modules separately at urllib — URL handling modules.
Install Beautiful Soup and choose a parser
The package is named beautifulsoup4 for installation and imported as bs4. The project’s PyPI page reports Beautiful Soup 4.15.0, released June 7, 2026, and a minimum Python version of 3.7; package metadata can change, so check the current PyPI project page if version compatibility is important.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors#1 Best Overall
python -m pip install beautifulsoup4
For the examples below, use Python’s built-in html.parser. Naming the parser explicitly makes the choice reproducible instead of depending on which optional parser happens to be installed. Beautiful Soup’s documentation describes these practical choices:
| Parser | When it can fit | Tradeoff |
|---|---|---|
html.parser |
Simple setup; it is included with Python and is documented as reasonably fast and lenient. | Malformed markup can produce a tree different from another parser’s tree. |
lxml HTML parser |
The documentation recommends it when speed matters and describes it as lenient. | It requires an external dependency; no numeric speed comparison is established here. |
html5lib |
Useful when browser-like HTML5 tree building is a priority; the documentation says it creates valid HTML5. | It requires an external dependency and is described as very slow. |
lxml XML parser |
Use for XML rather than HTML when you need the supported lxml XML parser. | Requires lxml; specify the XML mode rather than treating it as the default HTML parser. |
Install optional parser packages according to their current installation instructions. For malformed HTML, parser choice can affect the resulting tree. The Beautiful Soup documentation notes these differences and recommends specifying a parser, particularly when code is distributed.
Fetch a page and parse its HTML
This runnable example uses the standard-library urllib.request to retrieve a page and Beautiful Soup to parse the returned HTML. It prints the page title and a list of links found in the response. Replace the example URL with a page you are permitted to access.
from urllib.request import Request, urlopen
from bs4 import BeautifulSoup
url = "https://example.com/"
request = Request(url, headers={"User-Agent": "Python HTML parser example"})
with urlopen(request, timeout=20) as response:
html = response.read()
encoding = response.headers.get_content_charset() or "utf-8"
soup = BeautifulSoup(html, "html.parser", from_encoding=encoding)
print("Title:", soup.title.get_text(" ", strip=True) if soup.title else "(no title)")
for link in soup.find_all("a", href=True):
label = link.get_text(" ", strip=True)
href = link.get("href")
print(label, href)
The explicit timeout prevents the request from waiting indefinitely, and the response’s declared character set is used when available. Network access can still fail because of DNS, connectivity, TLS, server errors, or a site’s access controls; those are acquisition problems, not Beautiful Soup parsing errors. If you already have HTML—for example, from a saved file or another HTTP client—skip the request code and pass that content directly to BeautifulSoup.
Rank #2
Find elements and extract useful values
Use find() for one match
find() returns the first matching element, or None if no match exists. Always account for the missing-match case before accessing a result’s methods.
article = soup.find("article")
if article is not None:
print(article.get_text(" ", strip=True))
Use find_all() for repeated tags
find_all() returns all matching tags. Attributes can be passed as keyword arguments, such as class_ for the HTML class attribute.
headings = soup.find_all("h2", class_="story-title")
for heading in headings:
print(heading.get_text(" ", strip=True))
Use CSS selectors when they express the structure clearly
select() accepts CSS selectors and returns a list; select_one() returns the first match or None. For example, article h2 finds second-level headings nested inside article elements, while p.intro matches paragraphs with the class intro.
for heading in soup.select("article h2"):
print(heading.get_text(" ", strip=True))
first_intro = soup.select_one("p.intro")
if first_intro:
print(first_intro.get_text(" ", strip=True))
Extract text, links, and other attributes
Call get_text(" ", strip=True) to join text fragments with spaces and trim surrounding whitespace. For attributes, use get(); it safely returns None when the attribute is missing.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
for link in soup.select("a"):
text = link.get_text(" ", strip=True)
href = link.get("href")
if href is not None:
print({"text": text, "href": href})
For images, the source is commonly in src, but lazy-loaded pages may use another attribute or JavaScript to fill it later. Inspect the actual markup before assuming a field name. Beautiful Soup’s documented search, selector, and text methods are described in its API documentation.
Turn extracted content into structured data
Once each record has a stable container, search within that container rather than across the whole document. This avoids accidentally pairing a heading from one item with a link from another.
records = []
for card in soup.select("article.card"):
title_tag = card.select_one("h2")
link_tag = card.select_one("a[href]")
summary_tag = card.select_one("p.summary")
records.append({
"title": title_tag.get_text(" ", strip=True) if title_tag else None,
"url": link_tag.get("href") if link_tag else None,
"summary": summary_tag.get_text(" ", strip=True) if summary_tag else None,
})
for record in records:
print(record)
Use None or another deliberate missing-value representation when a field is absent. Avoid silently substituting unrelated values: a missing selector may signal changed markup, a different page variant, or content that is not in the HTML response.
Handle relative links and save results
Many pages use relative links such as /about or ../story. Resolve them against the page URL with Python’s urllib.parse.urljoin before saving or requesting them.
from urllib.parse import urljoin
page_url = "https://example.com/news/"
links = []
for tag in soup.select("a[href]"):
links.append({
"text": tag.get_text(" ", strip=True),
"url": urljoin(page_url, tag["href"]),
})
After extraction, the records are ordinary Python data and can be written as JSON or CSV using Python’s standard libraries. Validate fields and types before treating scraped values as trusted input; page content can be missing, inconsistent, or changed without notice.
What to check when a result is missing
- Inspect the received HTML. Confirm the target text or element is actually in the response body. The browser’s rendered view is not proof that the raw response contains it.
- Verify the selector against the markup. Check tag names, class names, nesting, and whether the HTML class attribute contains multiple classes.
- Handle absent matches. A
find()orselect_one()result can beNone; a selector returning no results is not itself a parser exception. - Try another parser when malformed markup is plausible. The same invalid HTML can yield different trees across parsers. Keep the parser explicit and compare the structure you actually get.
- Separate retrieval from parsing. If the response is an error page, challenge page, or incomplete body, changing a CSS selector will not recover the intended page content.
- Check for client-rendered content. If the value is absent from the HTML passed to Beautiful Soup, the parser cannot extract it. The official sources cited here do not establish how any particular JavaScript-heavy site behaves.
Be considerate when collecting web data
Check a site’s rules and requirements before crawling it. The Robots Exclusion Protocol, specified by IETF RFC 9309 (September 2022), defines rules that crawlers are requested to honor. A robots.txt file does not, by itself, settle every question about permission, contracts, or applicable law. Keep request volume reasonable and avoid collecting information you are not entitled to use.
Or skip the browser setup
If your goal is an image or PDF of a page rather than structured fields extracted from its HTML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. For example, save a WebP capture with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for request options. The capture can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchSign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
Best Value
Frequently Asked Questions
Can Beautiful Soup parse HTML I already saved to a file?
Yes. Read the file’s contents and pass the resulting text or bytes to BeautifulSoup with an explicit parser; no HTTP request is needed.
Does Beautiful Soup automatically follow links or crawl a website?
No. It parses one supplied document. Your code must separately decide whether and how to request linked URLs.
Can Beautiful Soup extract data from a PDF?
Beautiful Soup is for HTML and XML markup, not PDF text extraction. Use a PDF-specific tool if you need the document’s text.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

