Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Python’s built-in html.parser when you need dependency-free, event-driven processing. Use Beautiful Soup when you need to search and navigate a document tree; select its parser backend explicitly: lxml is very fast, while html5lib most closely follows browser-style recovery of broken HTML but is very slow. The right choice depends on whether your priority is zero dependencies, convenient extraction, speed, or tolerance of malformed markup.

This guide shows complete, reproducible examples for each approach, explains what parsing can and cannot do, and gives a troubleshooting path for real projects. Parsing begins only after HTML text or a file is available; downloading a URL and rendering JavaScript are separate concerns.

What HTML parsing does

An HTML parser turns markup text into events or a tree that your program can inspect. For example, the source <h1>News</h1><p>Read this</p> contains a heading element and a paragraph. A parser identifies those elements, their attributes, and their text so code can extract, transform, or validate content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python’s markup-processing modules include html.parser. The Python documentation describes an HTMLParser instance as being fed HTML data and calling handler methods for start tags, end tags, text, comments, and other markup. Beautiful Soup is a higher-level library that builds a navigable tree around a parser backend.

Choose a parser

Choice Best fit Trade-offs
html.parser Small scripts, standard-library deployments, handler-based processing No third-party install; event-oriented API; it parses invalid markup but does not validate matching tags.
Beautiful Soup + lxml Tree navigation when speed matters Requires the external lxml package and its C dependency.
Beautiful Soup + html5lib Browser-like recovery of badly formed HTML Very lenient and very slow; requires an external Python package.
Beautiful Soup + html.parser Convenient tree API without an additional parser package Recovery behavior differs from the other backends on malformed input.

Beautiful Soup’s documentation warns that different backends can create different trees for invalid markup. Therefore, pass the backend name explicitly in production code and tests.

Parse simple HTML with Python’s standard library

Collect text from selected tags

Subclass HTMLParser and implement the callbacks you need. This example collects heading and paragraph text while preserving the order in which it appears:

from html.parser import HTMLParser

class TextExtractor(HTMLParser):
    def __init__(self):
        super().__init__(convert_charrefs=True)
        self.capture = False
        self.current_tag = None
        self.parts = []
        self.records = []

    def handle_starttag(self, tag, attrs):
        if tag in {"h1", "h2", "p"}:
            self.capture = True
            self.current_tag = tag
            self.parts = []

    def handle_data(self, data):
        if self.capture:
            self.parts.append(data)

    def handle_endtag(self, tag):
        if self.capture and tag == self.current_tag:
            text = " ".join("".join(self.parts).split())
            if text:
                self.records.append((tag, text))
            self.capture = False
            self.current_tag = None
            self.parts = []

html = """

Parser guide

Use a handler for a small extraction task.

Details

""" parser = TextExtractor() parser.feed(html) parser.close() print(parser.records) # [('h1', 'Parser guide'), ('p', 'Use a handler for a small extraction task.'), # ('h2', 'Details')]

convert_charrefs=True is the documented Python 3.10 default. Character references are converted except in contexts such as script and style. Calling close() after the final feed() lets the parser finish buffered data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read attributes

The attrs argument is a list of (name, value) pairs. Convert it to a dictionary when duplicate attributes are not relevant to your task:

from html.parser import HTMLParser

class LinkParser(HTMLParser):
    def handle_starttag(self, tag, attrs):
        if tag == "a":
            attributes = dict(attrs)
            href = attributes.get("href")
            label = attributes.get("title", "(no title)")
            print(href, label)

parser = LinkParser()
parser.feed('Docs')
parser.close()

Keep the list form if duplicate attributes matter. HTMLParser is not a strict nesting validator: it does not check that end tags match start tags, and an element implicitly closed by an outer element may not trigger the corresponding end-tag callback. Treat it as a tolerant event stream, not as an HTML conformance checker.

Handle comments and declarations

from html.parser import HTMLParser

class MarkupEvents(HTMLParser):
    def handle_comment(self, data):
        print("comment:", data)

    def handle_decl(self, decl):
        print("declaration:", decl)

    def handle_startendtag(self, tag, attrs):
        print("self-closing:", tag, attrs)

p = MarkupEvents()
p.feed('<!DOCTYPE html><!-- note --><br/>')
p.close()

Parse HTML as a searchable tree with Beautiful Soup

Install and choose a backend

Install Beautiful Soup with the backend you intend to run:

python -m pip install beautifulsoup4
# Optional alternatives:
python -m pip install lxml html5lib

Pass "html.parser", "lxml", or "html5lib" explicitly. Omitting the backend can make behavior depend on which packages happen to be installed.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract headings, links and visible text

from bs4 import BeautifulSoup

html = """

  

Parser guide

Choose a backend explicitly.

Documentation """ soup = BeautifulSoup(html, "html.parser") print(soup.h1.get_text(" ", strip=True)) for link in soup.select("a[href]"): print(link.get("href"), link.get_text(" ", strip=True)) print(soup.get_text(" ", strip=True))

select() accepts CSS selectors, while methods such as find(), find_all(), and attribute access support simpler searches. get_text(" ", strip=True) inserts a separator so words from adjacent elements do not run together.

Modify or remove nodes

from bs4 import BeautifulSoup

soup = BeautifulSoup('

Keep

Remove
', "html.parser") for ad in soup.select(".ad"): ad.decompose() soup.main.append(soup.new_tag("p")) soup.main.p.string = "Added by the parser" print(soup.main.decode())

Beautiful Soup converts input to Unicode and accepts either markup text or an open file handle. Its tree model is convenient for edits, but the exact tree for malformed input is backend-dependent.

Backend behavior and malformed markup

html.parser

Use it when a standard-library dependency and direct callbacks are more important than a high-level query API. It accepts imperfect markup, but its event behavior is not a browser’s tree-building algorithm and it does not verify matching start and end tags.

lxml

Use Beautiful Soup with lxml when you want the same convenient tree interface with the speed-oriented backend described in the documentation. Plan for an external C dependency in your deployment and lock the package versions used by your tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

html5lib

Choose html5lib when browser-like HTML5 recovery is more important than runtime. The documentation characterizes it as extremely lenient and very slow, and it requires an external Python package.

Why explicit selection matters

Malformed source can produce different parent-child relationships under each backend. A selector that succeeds with one tree can fail with another. Set the backend in every BeautifulSoup(...) call, add malformed fixtures to tests, and compare the resulting tree when upgrading dependencies.

Parse a file safely

Open files with the encoding your producer specifies. If the encoding is unknown, determine it from reliable metadata rather than silently assuming one:

from pathlib import Path
from bs4 import BeautifulSoup

source = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(source, "html.parser")
title = soup.title.get_text(strip=True) if soup.title else None
print(title)

For large documents, avoid retaining unnecessary trees or text. A streaming HTMLParser handler can emit records as events arrive; Beautiful Soup generally builds a complete in-memory tree, which is easier to query but consumes more memory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parsing is not downloading or JavaScript rendering

These parsers operate on HTML you already have. Getting a response from a remote URL introduces separate questions: HTTP status handling, timeouts, redirects, response encoding, authentication, and robots or access policies. The parser documentation does not define a complete network-fetching workflow, so choose and document an HTTP client independently.

Likewise, parsing the original response does not execute JavaScript. If content is inserted after page load, obtain the rendered HTML with an appropriate browser automation workflow, then pass that resulting HTML to your parser. Do not assume that a parser can reveal data absent from its input.

Or skip the browser setup

If your goal is to obtain a clean snapshot of a URL before parsing, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

cURL example (see the ScreenshotNeo API documentation):

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server so Claude, Cursor, and other MCP clients can call take_screenshot, get_page_info, and capture_pdf. Free accounts include 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

“Feature X is missing”

Check which object you created. HTMLParser exposes callbacks, not Beautiful Soup’s select() or find_all(). Either implement the needed callback or parse the same input with Beautiful Soup.

Different output on another machine

A backend may differ or be selected implicitly. Install the intended backend, pass its name explicitly, and pin compatible dependency versions.

Text is duplicated or runs together

HTML often contains nested formatting elements and whitespace nodes. Use get_text(" ", strip=True) in Beautiful Soup, or accumulate and normalize data in your handle_data implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Expected closing-tag code never runs

HTMLParser does not call an end-tag handler for every implicit close and does not validate nesting. Track only the state your extraction requires, or use a tree parser when parent-child relationships are essential.

Expected content is absent

Print or save the exact HTML passed to the parser. The content may be loaded by JavaScript, behind authentication, or absent from the response entirely; parsing cannot manufacture it.

Non-ASCII characters are corrupted

Decode bytes with the producer’s declared encoding before parsing, and use an explicit encoding when reading files. Do not “fix” mojibake after the tree has already been built.

Performance, reliability and maintenance

  • For small, predictable extraction jobs, html.parser minimizes installation and startup overhead.
  • For many CSS-style queries or edits, Beautiful Soup reduces custom state-machine code; choose lxml when its external dependency is acceptable and speed is a priority.
  • Use html5lib only when its browser-like recovery justifies very slow processing.
  • Keep parser selection, input encoding, and extraction rules in tests. Include valid and malformed fixtures because backend recovery can change selectors.
  • Limit retained nodes and stream with callbacks when document size makes a complete tree impractical.
  • Separate acquisition from parsing so HTTP failures, rendering failures, and extraction failures can be diagnosed independently.

Frequently Asked Questions

Can HTMLParser validate that HTML is correctly nested?

No. Python’s HTMLParser is designed for event handling and does not check whether end tags match start tags. Use a dedicated validator when conformance checking is required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which Beautiful Soup parser should I use for reproducible tests?

Choose one explicitly in every call and pin the corresponding dependency versions. The best backend depends on whether your project values standard-library availability, speed, or browser-like recovery.

Can these parsers execute JavaScript?

No. They parse supplied HTML only. Obtain rendered markup separately, then parse that result.

The Bottom Line

Start with html.parser for dependency-free callbacks; choose Beautiful Soup for tree navigation, with an explicitly selected backend. Use lxml for speed when its dependency is acceptable and html5lib for browser-like recovery when performance is secondary.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.