Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
Beautiful Soup

How to Parse HTML in Python: A Step-by-Step Guide for Beginners

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To parse HTML in Python, start with markup you already have, then turn it into a structure your code can inspect. For a beginner-friendly tree interface, use Beautiful Soup: create a BeautifulSoup object, choose a parser explicitly, and find the elements, text, or attributes you need. Python’s built-in html.parser is another option for event-driven tasks; lxml is useful when its HTML or XML APIs fit your input. Parsing is not the same as downloading a page or running its JavaScript.

How do I parse HTML in Python?

Parsing means interpreting markup and organizing it into objects or events your program can inspect. It starts after you have the HTML as a string or file. It does not, by itself, make an HTTP request, render JavaScript, or give permission to collect information from a website.

For most beginner tasks that involve searching nested content, Beautiful Soup provides a convenient tree. The example below parses a string, finds a heading, extracts a link, and prints the text:

from bs4 import BeautifulSoup

html = """
<article>
  <h1>A Python example</h1>
  <p>Read the <a href="/guide">guide</a>.</p>
</article>
"""

soup = BeautifulSoup(html, "html.parser")

heading = soup.find("h1")
link = soup.find("a")

print(heading.get_text())          # A Python example
print(link.get_text())             # guide
print(link.get("href"))            # /guide

Beautiful Soup is a third-party package. Install it in the Python environment where you run your script with python -m pip install beautifulsoup4. The parser name in the second argument, here html.parser, tells Beautiful Soup how to interpret the markup. Supplying that choice explicitly makes the intended parser clear and helps avoid environment-dependent results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse HTML from a file

Read the file, then pass its contents to the same constructor. Use an explicit encoding when you know the file’s character encoding; UTF-8 is common, but not guaranteed for every source.

from pathlib import Path
from bs4 import BeautifulSoup

html = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")

for paragraph in soup.find_all("p"):
    print(paragraph.get_text(" ", strip=True))

Beautiful Soup also accepts a file-like object. If a file’s encoding is uncertain, determine it from reliable information about that file rather than assuming that every input is UTF-8.

Install a different parser when needed

Beautiful Soup is an interface over a parser; it does not make the parser choices interchangeable in every detail. The named parsers described in its documentation include Python’s built-in html.parser, lxml, and html5lib. To use an optional parser, install its corresponding package in your environment. The exact package and installation method depend on the parser you select.

# Examples of explicit parser selection after installing the relevant parser:
soup = BeautifulSoup(html, "lxml")
# or
soup = BeautifulSoup(html, "html5lib")

How do I extract text from HTML in Python?

Find the element or elements of interest, then request their text. get_text() returns text inside a tag, including text in nested tags. Its optional separator can keep adjacent text nodes from running together; strip=True removes whitespace at the ends.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

html = "<div><p>One <em>important</em> point.</p><p>Next.</p></div>"
soup = BeautifulSoup(html, "html.parser")

paragraphs = [
    p.get_text(" ", strip=True)
    for p in soup.find_all("p")
]
print(paragraphs)
# ['One important point.', 'Next.']

For a single matching element, find() returns the first match or None if no match exists. For multiple matches, find_all() returns a collection you can iterate over. Check for a missing element before calling methods on it:

title = soup.find("h1")
if title is None:
    print("No h1 element found")
else:
    print(title.get_text(" ", strip=True))

Text extraction is not the same as preserving the document’s visual layout. A parser reads markup, not the browser’s rendered appearance; whitespace and line breaks in extracted text may differ from what a person sees on a page.

How do I find HTML elements and attributes?

Use a tag name, an attribute, or a CSS selector to narrow a search. Attributes such as href and class are available on the parsed tag. An absent attribute returns None with get(), so handle that case if it matters to your program.

html = """
<ul>
  <li class="result"><a href="/one">First</a></li>
  <li class="result"><a href="/two">Second</a></li>
</ul>
"""
soup = BeautifulSoup(html, "html.parser")

for item in soup.select("li.result"):
    anchor = item.find("a")
    if anchor is not None:
        print(anchor.get_text(" ", strip=True), anchor.get("href"))

select() accepts CSS selectors, which can be handy when an element is identified by a class or relationship. find() and find_all() are direct choices for searching by tag and other supported criteria. Prefer selectors that reflect meaningful structure rather than assuming an incidental class name will never change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which Python HTML parser should I choose?

Option Best fit Trade-offs
html.parser A small task that fits the standard library, especially when callbacks should react as markup is read. You implement event handling by subclassing HTMLParser. Its documentation says it does not check that end tags match start tags.
Beautiful Soup A Python-friendly tree for finding, reading, and navigating nested elements. It is an interface over a selected parser. Parser choice can affect the resulting tree, especially for malformed markup; specify the parser for repeatability.
lxml Tasks that suit its HTML or XML parsing APIs, including XHTML when XML rules are intended. Be deliberate about whether the input should be treated as HTML or XML. Parsing XHTML as HTML can produce unexpected results when XML semantics are intended.

There is no universal performance winner established here: a comparable, task-specific benchmark is not available. Choose based on whether you need callbacks or a tree, how you want to handle imperfect markup, which dependencies are available in your deployment environment, and whether the input is HTML or XHTML/XML.

Use Python’s built-in event-driven parser

The standard-library html.parser module lets you subclass HTMLParser and override handlers for start tags, end tags, and data. Python’s documentation describes an instance being fed HTML and calling handler methods as markup elements are encountered. This callback model suits processing that can happen as events arrive, but it does not automatically give you the same convenient search-and-navigation workflow as a tree.

from html.parser import HTMLParser

class TextCollector(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []

    def handle_data(self, data):
        text = data.strip()
        if text:
            self.parts.append(text)

parser = TextCollector()
parser.feed("<p>Hello <strong>Python</strong>.</p>")
print(" ".join(parser.parts))  # Hello Python .

This minimal collector deliberately captures text events, not a nested document tree or carefully formatted prose. For richer callback behavior, implement the handlers appropriate to the task. The parser’s documented lack of start/end tag matching checks means you should not treat successful event handling as validation that the markup is well-formed.

What changes when HTML is malformed?

Real-world markup can contain missing, mismatched, or unexpectedly nested tags. Different parsers may construct different trees from the same malformed input. As a result, an element can appear under a different parent than expected, or a search can return a different result than it does with another parser.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Set the parser explicitly instead of relying on whichever parser happens to be installed.
  • When an expected element is missing, inspect the parsed structure and the original input around that element.
  • Test against representative input, including imperfect cases, rather than assuming the source markup is always valid.
  • If code runs on multiple machines, ensure the selected parser is available in each environment.

How should I handle XHTML?

XHTML is XML-based markup. If the input is XHTML and you need XML rules, parse it as XML rather than assuming HTML parsing will preserve the intended semantics. The lxml project specifically advises that XHTML is preferably parsed as XML; treating it as HTML can lead to unexpected results. Make the format decision based on what the input actually is and what rules the task requires.

What parsing does not do: fetching and rendering a page

A parser consumes markup you supply. It does not retrieve a URL, execute client-side JavaScript, or wait for a browser-rendered page. If you need to obtain HTML, fetching is a separate step; if the content is created only after JavaScript runs, parsing a static response may not contain it. Keep those tasks separate so you can diagnose whether an issue comes from acquisition, rendering, or parsing. Follow the site’s access rules and applicable requirements when obtaining web content.

Or skip the browser setup

If what you need is a screenshot or PDF rather than a parsed HTML tree, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It does not replace Beautiful Soup for extracting structured text or attributes. One GET request captures a URL; see the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo free.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common parsing problems

  • ModuleNotFoundError: No module named 'bs4': Beautiful Soup is not installed in the Python environment running the script. Install beautifulsoup4 with that environment’s Python, for example python -m pip install beautifulsoup4.
  • An optional parser cannot be found: You selected a parser such as lxml or html5lib without making it available in the active environment. Install the parser you intend to use, or select html.parser.
  • AttributeError after a search: A find() call may have returned None because the tag is absent or the parsed structure differs from your expectation. Check the result before calling .get_text() or accessing its attributes.
  • Unexpected nesting or missing elements: The source may be malformed, or another parser may build a different tree. Specify the parser and inspect the parsed output around the affected content.
  • Text looks concatenated or oddly spaced: Nested tags create separate text pieces. Try get_text(" ", strip=True), then adjust whitespace handling for the output you actually need.
  • Expected content is not in the input: Parsing cannot add content absent from the supplied markup. Confirm how the HTML was obtained and whether the content depends on browser-side JavaScript before changing parsing code.

Further reading

If you want to progress from this beginner task into broader scraping techniques, O’Reilly lists Ryan Mitchell’s Web Scraping with Python, 3rd Edition, published in February 2024, and describes it as intermediate to advanced. It is optional further reading, not a prerequisite for parsing a string or file.

Frequently Asked Questions

Can Python parse HTML without installing a package?

Yes. The standard-library html.parser module is available without Beautiful Soup; it uses event handlers rather than Beautiful Soup’s tree-search interface.

Does Beautiful Soup download a web page for me?

No. It parses markup passed to it. Retrieving a URL is a separate operation.

Which parser should a beginner start with?

For searching and navigating nested content, Beautiful Soup with an explicitly selected parser is a practical starting point. Use the standard library when callbacks fit your task.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.