Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For most straightforward HTML extraction, start with Beautiful Soup; for direct tree work where response time matters, consider lxml; choose html5lib when browser-aligned HTML5 parsing rules matter most; use Python’s built-in html.parser when avoiding an extra dependency is the priority; and benchmark selectolax when CSS selectors and throughput are central to your workload. There is no universally best parser: the right choice depends on how much malformed HTML you receive, what tree you need, and whether ease, standards behavior, dependencies, or performance matters most.

What counts as a Python HTML parser?

“HTML parser” can mean the engine that turns markup into a tree, or the Python interface you use to search that tree and extract data. That distinction is important for Beautiful Soup: it provides a consistent, approachable Python-facing API, but delegates parsing to a backend such as Python’s html.parser, lxml, or html5lib. Changing that backend can change both speed and the tree produced from imperfect markup.

These five choices serve different needs. This is a practical shortlist, not a universal performance ranking. In particular, HTML written by people or generated by websites is often malformed or incomplete; parsers may repair it in different ways, so choose according to the rules your application needs rather than assuming all libraries produce an identical document tree.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At a glance: the five libraries

Library Best fit Main trade-off
Beautiful Soup Readable, approachable extraction code; convenient searching across documents It wraps a backend, whose availability and behavior affect results and speed
lxml Direct HTML or XML tree work, especially when response time matters Check that its handling of malformed input matches your needs
html5lib Parsing behavior designed to conform to the WHATWG HTML specification Standards-oriented parsing can come with a speed trade-off; no universal ratio is established here
html.parser A built-in starting point with no separate parser package Its resulting tree can differ from other parsers, especially on malformed HTML
selectolax CSS-selector extraction and a candidate to benchmark for throughput Published benchmark results are project-produced and workload-specific

1. Beautiful Soup: the approachable extraction interface

Beautiful Soup is a strong starting point when your main goal is to write clear extraction code rather than work with a parser’s lower-level interfaces. It supports familiar operations such as finding elements by tag, filtering by attributes, and reading text. Crucially, Beautiful Soup is not one fixed parsing engine: you select a backend, and the backend influences how markup is interpreted.

Make the backend explicit

Beautiful Soup’s documentation says its default is the best parser installed. That means two machines can silently choose different backends if their installed dependencies differ. When reproducible behavior matters, specify the backend in the code and keep the environment’s dependencies consistent.

from bs4 import BeautifulSoup

html = """
<html><body>
  <article><h1>Example title</h1><a href="/story">Read</a></article>
</body></html>
"""
soup = BeautifulSoup(html, "lxml")

print(soup.select_one("article h1").get_text(strip=True))
print(soup.select_one("article a")["href"])

Use another installed backend by changing the second argument, for example to "html.parser" or "html5lib". If the application depends on exactly how unusual input is repaired, pin that choice rather than relying on whichever parser happens to be installed.

When it is the right choice

  • Choose it when readable extraction code and an easy search interface are more important than direct access to a particular parser’s tree API.
  • Choose it when you want the option to change parser backends without rewriting all of your extraction logic.
  • If response time is critical, Beautiful Soup’s own documentation advises working directly with lxml; it also says Beautiful Soup parses significantly faster with lxml than with html.parser or html5lib.

2. lxml: direct HTML and XML tree work

lxml is the direct choice to evaluate when you need HTML or XML tree facilities and care about performance. You can parse HTML and use XPath to find nodes without adding Beautiful Soup’s higher-level interface. The Beautiful Soup documentation specifically directs users with critical response-time needs to work directly atop lxml rather than expect the wrapper to match the speed of its underlying parser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from lxml import html

markup = """
<article><h1>Example title</h1><a href="/story">Read</a></article>
"""
tree = html.fromstring(markup)

title = tree.xpath("string(//article/h1)").strip()
link = tree.xpath("string(//article/a/@href)")
print(title)
print(link)

Direct lxml is a good fit if XPath and direct tree work suit your application. It is also an option underneath Beautiful Soup: choosing "lxml" as the backend can preserve Beautiful Soup’s extraction interface while changing the parser engine.

Do not select it by speed alone

When markup is malformed, compare the resulting tree with what the application expects. Faster processing is not useful if a parser’s repairs cause you to extract the wrong node or miss content. If the input is irregular, build a small regression sample from real pages and assert the fields your application needs.

3. html5lib: when HTML5 parsing rules matter

html5lib’s project describes it as designed to conform to the WHATWG HTML specification as implemented by major web browsers. That makes it a candidate when standards-oriented handling of real-world HTML is more important than raw speed. This description is the project’s stated design goal, not an independent conformance audit.

import html5lib

markup = "<article><h1>Example title</h1></article>"
document = html5lib.parse(markup)

# html5lib's default tree builder returns an ElementTree-like document.
root = document.getroot()
for element in root.iter():
    if element.tag.endswith("h1"):
        print("".join(element.itertext()).strip())

html5lib can use different tree builders, including ElementTree, minidom, and lxml.etree. Choose the output representation that fits the rest of your code, and verify the API details against the version installed in your project.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The trade-off

Standards-oriented parsing can be a poor fit for high-volume jobs if speed is the overriding constraint. The material available for this comparison does not establish a general speed ratio between html5lib and every other parser, so benchmark your own pages and extraction work rather than rely on a blanket claim.

4. Python’s built-in html.parser: no extra parser package

html.parser is included in Python’s standard library. It is a sensible first choice when you want to avoid an additional parsing dependency and can handle extraction through its event-style interface. Unlike Beautiful Soup’s search-oriented API, you typically subclass HTMLParser and define what to do when tags and text arrive.

from html.parser import HTMLParser

class TitleParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_title = False
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag == "h1":
            self.in_title = True

    def handle_endtag(self, tag):
        if tag == "h1":
            self.in_title = False

    def handle_data(self, data):
        if self.in_title:
            self.parts.append(data)

parser = TitleParser()
parser.feed("<article><h1>Example title</h1></article>")
print("".join(parser.parts).strip())

This can be sufficient for narrow tasks, especially when you need to react to a small set of tags as input is parsed. If you want to query a rich tree repeatedly, another library may offer a more convenient fit. And if the input includes malformed markup, check the output rather than assume the same repairs as a browser or another library.

5. selectolax: CSS selectors and a throughput candidate

selectolax provides HTML parsing and CSS-selector queries. Its project recommends the Lexbor backend and documents a CSS-selector workflow using LexborHTMLParser and css_first. It is worth benchmarking when your extraction consists of selector-based lookups and throughput matters.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from selectolax.lexbor import LexborHTMLParser

markup = """
<article><h1>Example title</h1><a href="/story">Read</a></article>
"""
tree = LexborHTMLParser(markup)
title = tree.css_first("article h1")
link = tree.css_first("article a")

print(title.text(strip=True) if title else None)
print(link.attributes.get("href") if link else None)

Check the project’s current installation and API guidance when setting up an environment: package interfaces and recommendations can change. The evidence for its speed is a project benchmark on a particular extraction task, not a neutral test of every library against every workload.

Why malformed HTML changes the answer

Parser choice can affect more than performance. Beautiful Soup’s documentation demonstrates the input <a></p> and shows distinct results: lxml drops the unmatched closing </p> and adds html and body; html5lib constructs a paragraph and adds html, head, and body; and html.parser leaves a simpler tree. For invalid input, there is no universally correct result unless you first specify the parsing behavior you want.

If extraction unexpectedly misses a field, inspect the parsed tree before changing selectors. For Beautiful Soup, its diagnose() helper reports how different parsers handle a piece of input:

from bs4 import BeautifulSoup

markup = "<a></p>"
BeautifulSoup.diagnose(markup)

Use this sort of comparison to determine whether the problem is an incorrect selector, a parser’s repair of invalid markup, or an assumption in your extraction code. Once you settle on the intended tree behavior, specify the backend explicitly and retain representative malformed examples in your tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the available performance numbers do—and do not—show

selectolax’s repository describes a simple benchmark that extracts titles, links, scripts, and a meta tag from the main pages of 754 domains. The repository reports these elapsed times for that task:

Configuration in the project benchmark Reported time
Beautiful Soup with html.parser 61.02 seconds
lxml / Beautiful Soup with the lxml backend 9.09 seconds
html5_parser 16.10 seconds
selectolax (Modest) 2.94 seconds
selectolax (Lexbor) 2.39 seconds

These are results from the selectolax project’s sample benchmark; the accessed repository material does not state a publication year for these figures. They describe one extraction task over those pages, not a general speed guarantee or a vendor-neutral comparison. Your HTML, selector patterns, machine, and output requirements may change the result. The same project benchmark favors Lexbor among the listed configurations, but that does not make it the fastest choice for every application.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to choose for your workload

  1. Start with the output you need. For a few clear extractions and readable code, try Beautiful Soup. For XPath or direct HTML/XML tree work, try lxml. For event-style processing without an extra dependency, consider html.parser.
  2. Decide what “correct” means for imperfect input. If WHATWG-aligned HTML parsing behavior matters, evaluate html5lib. If a particular malformed page produces a surprising result, inspect and compare the parsed trees rather than expecting every parser to agree.
  3. Make repeatability explicit. If you use Beautiful Soup, select a backend in code instead of accepting the best parser installed on each machine.
  4. Measure the task you actually run. If response time or throughput determines the choice, use representative HTML and the same extraction logic to compare the candidates. Do not infer your job’s performance from a benchmark with a different workload.
  5. Test the edge cases that matter. Include missing tags, mismatched closing tags, empty content, and the selector variations your input can contain. Confirm that your extraction handles absent nodes instead of assuming every lookup succeeds.

Where ScreenshotNeo fits—and where it does not

ScreenshotNeo is a website screenshot API and MCP server, not an HTML parsing library or a replacement for Beautiful Soup, lxml, html5lib, html.parser, or selectolax. Use a parser when you need to inspect markup as a tree. If the task is instead to capture a website as an image or PDF, ScreenshotNeo is an alternative to try first: a single request can return a screenshot, and its response reports whether a capture was clean, failed, or a cache hit. Its clean-capture options accept cookie or consent banners and remove supported consent platforms, newsletter popups, and chat widgets before the shot.

Or skip the browser setup

Use ScreenshotNeo for a screenshot rather than a parsed DOM tree. This cURL example captures a page as WebP; replace the sample URL as needed. See the ScreenshotNeo API documentation for request options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. If a rendered capture is what you need, sign up for ScreenshotNeo’s free plan.

Common problems and fixes

  • The parsed tree differs between machines. With Beautiful Soup, different installed backends can change the default parser. Specify the backend, keep dependencies aligned, and test the chosen behavior.
  • A selector finds nothing on malformed markup. Inspect the parsed tree and compare backend behavior. Use Beautiful Soup’s diagnose() helper when applicable, then choose parsing rules that match your intended result.
  • Beautiful Soup is too slow for a critical path. Its documentation says it will not be as fast as the parsers it sits on top of. Try its lxml backend first; for response-time-critical code, evaluate using lxml directly.
  • The result omits content added by JavaScript. A parser does not render JavaScript. Parsing HTML source alone does not establish what a browser would show after scripts run; use a rendering or capture workflow if the rendered page is the required input.
  • A benchmark winner is slower on your job. The published selectolax figures cover a specific task. Benchmark representative input and identical extraction work in your own environment before choosing on speed.

Frequently asked questions

Is Beautiful Soup itself an HTML parser?

It is a Python-facing parsing and extraction interface that delegates markup parsing to a selected backend. The backend affects the resulting tree.

Which option follows browser-style HTML parsing rules?

html5lib is designed to conform to the WHATWG HTML specification. If that behavior is essential, evaluate it with your input and chosen tree builder.

Can a Python HTML parser retrieve content created by JavaScript?

No. Parsing markup and rendering a page in a browser are separate tasks; a parser alone does not execute page scripts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.