Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For most straightforward HTML extraction, start with Beautiful Soup; for direct tree work where response time matters, consider lxml; choose html5lib when browser-aligned HTML5 parsing rules matter most; use Python’s built-in html.parser when avoiding an extra dependency is the priority; and benchmark selectolax when CSS selectors and throughput are central to your workload. There is no universally best parser: the right choice depends on how much malformed HTML you receive, what tree you need, and whether ease, standards behavior, dependencies, or performance matters most.
What counts as a Python HTML parser?
“HTML parser” can mean the engine that turns markup into a tree, or the Python interface you use to search that tree and extract data. That distinction is important for Beautiful Soup: it provides a consistent, approachable Python-facing API, but delegates parsing to a backend such as Python’s html.parser, lxml, or html5lib. Changing that backend can change both speed and the tree produced from imperfect markup.
These five choices serve different needs. This is a practical shortlist, not a universal performance ranking. In particular, HTML written by people or generated by websites is often malformed or incomplete; parsers may repair it in different ways, so choose according to the rules your application needs rather than assuming all libraries produce an identical document tree.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →At a glance: the five libraries
| Library | Best fit | Main trade-off |
|---|---|---|
| Beautiful Soup | Readable, approachable extraction code; convenient searching across documents | It wraps a backend, whose availability and behavior affect results and speed |
| lxml | Direct HTML or XML tree work, especially when response time matters | Check that its handling of malformed input matches your needs |
| html5lib | Parsing behavior designed to conform to the WHATWG HTML specification | Standards-oriented parsing can come with a speed trade-off; no universal ratio is established here |
html.parser |
A built-in starting point with no separate parser package | Its resulting tree can differ from other parsers, especially on malformed HTML |
| selectolax | CSS-selector extraction and a candidate to benchmark for throughput | Published benchmark results are project-produced and workload-specific |
1. Beautiful Soup: the approachable extraction interface
Beautiful Soup is a strong starting point when your main goal is to write clear extraction code rather than work with a parser’s lower-level interfaces. It supports familiar operations such as finding elements by tag, filtering by attributes, and reading text. Crucially, Beautiful Soup is not one fixed parsing engine: you select a backend, and the backend influences how markup is interpreted.
#1 Best Overall
Make the backend explicit
Beautiful Soup’s documentation says its default is the best parser installed. That means two machines can silently choose different backends if their installed dependencies differ. When reproducible behavior matters, specify the backend in the code and keep the environment’s dependencies consistent.
from bs4 import BeautifulSoup
html = """
<html><body>
<article><h1>Example title</h1><a href="/story">Read</a></article>
</body></html>
"""
soup = BeautifulSoup(html, "lxml")
print(soup.select_one("article h1").get_text(strip=True))
print(soup.select_one("article a")["href"])
Use another installed backend by changing the second argument, for example to "html.parser" or "html5lib". If the application depends on exactly how unusual input is repaired, pin that choice rather than relying on whichever parser happens to be installed.
When it is the right choice
- Choose it when readable extraction code and an easy search interface are more important than direct access to a particular parser’s tree API.
- Choose it when you want the option to change parser backends without rewriting all of your extraction logic.
- If response time is critical, Beautiful Soup’s own documentation advises working directly with lxml; it also says Beautiful Soup parses significantly faster with lxml than with
html.parseror html5lib.
2. lxml: direct HTML and XML tree work
lxml is the direct choice to evaluate when you need HTML or XML tree facilities and care about performance. You can parse HTML and use XPath to find nodes without adding Beautiful Soup’s higher-level interface. The Beautiful Soup documentation specifically directs users with critical response-time needs to work directly atop lxml rather than expect the wrapper to match the speed of its underlying parser.
from lxml import html
markup = """
<article><h1>Example title</h1><a href="/story">Read</a></article>
"""
tree = html.fromstring(markup)
title = tree.xpath("string(//article/h1)").strip()
link = tree.xpath("string(//article/a/@href)")
print(title)
print(link)
Direct lxml is a good fit if XPath and direct tree work suit your application. It is also an option underneath Beautiful Soup: choosing "lxml" as the backend can preserve Beautiful Soup’s extraction interface while changing the parser engine.
Rank #2
Do not select it by speed alone
When markup is malformed, compare the resulting tree with what the application expects. Faster processing is not useful if a parser’s repairs cause you to extract the wrong node or miss content. If the input is irregular, build a small regression sample from real pages and assert the fields your application needs.
3. html5lib: when HTML5 parsing rules matter
html5lib’s project describes it as designed to conform to the WHATWG HTML specification as implemented by major web browsers. That makes it a candidate when standards-oriented handling of real-world HTML is more important than raw speed. This description is the project’s stated design goal, not an independent conformance audit.
import html5lib
markup = "<article><h1>Example title</h1></article>"
document = html5lib.parse(markup)
# html5lib's default tree builder returns an ElementTree-like document.
root = document.getroot()
for element in root.iter():
if element.tag.endswith("h1"):
print("".join(element.itertext()).strip())
html5lib can use different tree builders, including ElementTree, minidom, and lxml.etree. Choose the output representation that fits the rest of your code, and verify the API details against the version installed in your project.
The trade-off
Standards-oriented parsing can be a poor fit for high-volume jobs if speed is the overriding constraint. The material available for this comparison does not establish a general speed ratio between html5lib and every other parser, so benchmark your own pages and extraction work rather than rely on a blanket claim.
4. Python’s built-in html.parser: no extra parser package
html.parser is included in Python’s standard library. It is a sensible first choice when you want to avoid an additional parsing dependency and can handle extraction through its event-style interface. Unlike Beautiful Soup’s search-oriented API, you typically subclass HTMLParser and define what to do when tags and text arrive.
from html.parser import HTMLParser
class TitleParser(HTMLParser):
def __init__(self):
super().__init__()
self.in_title = False
self.parts = []
def handle_starttag(self, tag, attrs):
if tag == "h1":
self.in_title = True
def handle_endtag(self, tag):
if tag == "h1":
self.in_title = False
def handle_data(self, data):
if self.in_title:
self.parts.append(data)
parser = TitleParser()
parser.feed("<article><h1>Example title</h1></article>")
print("".join(parser.parts).strip())
This can be sufficient for narrow tasks, especially when you need to react to a small set of tags as input is parsed. If you want to query a rich tree repeatedly, another library may offer a more convenient fit. And if the input includes malformed markup, check the output rather than assume the same repairs as a browser or another library.
5. selectolax: CSS selectors and a throughput candidate
selectolax provides HTML parsing and CSS-selector queries. Its project recommends the Lexbor backend and documents a CSS-selector workflow using LexborHTMLParser and css_first. It is worth benchmarking when your extraction consists of selector-based lookups and throughput matters.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
from selectolax.lexbor import LexborHTMLParser
markup = """
<article><h1>Example title</h1><a href="/story">Read</a></article>
"""
tree = LexborHTMLParser(markup)
title = tree.css_first("article h1")
link = tree.css_first("article a")
print(title.text(strip=True) if title else None)
print(link.attributes.get("href") if link else None)
Check the project’s current installation and API guidance when setting up an environment: package interfaces and recommendations can change. The evidence for its speed is a project benchmark on a particular extraction task, not a neutral test of every library against every workload.
Why malformed HTML changes the answer
Parser choice can affect more than performance. Beautiful Soup’s documentation demonstrates the input <a></p> and shows distinct results: lxml drops the unmatched closing </p> and adds html and body; html5lib constructs a paragraph and adds html, head, and body; and html.parser leaves a simpler tree. For invalid input, there is no universally correct result unless you first specify the parsing behavior you want.
If extraction unexpectedly misses a field, inspect the parsed tree before changing selectors. For Beautiful Soup, its diagnose() helper reports how different parsers handle a piece of input:
from bs4 import BeautifulSoup
markup = "<a></p>"
BeautifulSoup.diagnose(markup)
Use this sort of comparison to determine whether the problem is an incorrect selector, a parser’s repair of invalid markup, or an assumption in your extraction code. Once you settle on the intended tree behavior, specify the backend explicitly and retain representative malformed examples in your tests.
Recommended Free Tools
What the available performance numbers do—and do not—show
selectolax’s repository describes a simple benchmark that extracts titles, links, scripts, and a meta tag from the main pages of 754 domains. The repository reports these elapsed times for that task:
Best Value
| Configuration in the project benchmark | Reported time |
|---|---|
Beautiful Soup with html.parser |
61.02 seconds |
| lxml / Beautiful Soup with the lxml backend | 9.09 seconds |
| html5_parser | 16.10 seconds |
| selectolax (Modest) | 2.94 seconds |
| selectolax (Lexbor) | 2.39 seconds |
These are results from the selectolax project’s sample benchmark; the accessed repository material does not state a publication year for these figures. They describe one extraction task over those pages, not a general speed guarantee or a vendor-neutral comparison. Your HTML, selector patterns, machine, and output requirements may change the result. The same project benchmark favors Lexbor among the listed configurations, but that does not make it the fastest choice for every application.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to choose for your workload
- Start with the output you need. For a few clear extractions and readable code, try Beautiful Soup. For XPath or direct HTML/XML tree work, try lxml. For event-style processing without an extra dependency, consider
html.parser. - Decide what “correct” means for imperfect input. If WHATWG-aligned HTML parsing behavior matters, evaluate html5lib. If a particular malformed page produces a surprising result, inspect and compare the parsed trees rather than expecting every parser to agree.
- Make repeatability explicit. If you use Beautiful Soup, select a backend in code instead of accepting the best parser installed on each machine.
- Measure the task you actually run. If response time or throughput determines the choice, use representative HTML and the same extraction logic to compare the candidates. Do not infer your job’s performance from a benchmark with a different workload.
- Test the edge cases that matter. Include missing tags, mismatched closing tags, empty content, and the selector variations your input can contain. Confirm that your extraction handles absent nodes instead of assuming every lookup succeeds.
Where ScreenshotNeo fits—and where it does not
ScreenshotNeo is a website screenshot API and MCP server, not an HTML parsing library or a replacement for Beautiful Soup, lxml, html5lib, html.parser, or selectolax. Use a parser when you need to inspect markup as a tree. If the task is instead to capture a website as an image or PDF, ScreenshotNeo is an alternative to try first: a single request can return a screenshot, and its response reports whether a capture was clean, failed, or a cache hit. Its clean-capture options accept cookie or consent banners and remove supported consent platforms, newsletter popups, and chat widgets before the shot.
Or skip the browser setup
Use ScreenshotNeo for a screenshot rather than a parsed DOM tree. This cURL example captures a page as WebP; replace the sample URL as needed. See the ScreenshotNeo API documentation for request options.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. If a rendered capture is what you need, sign up for ScreenshotNeo’s free plan.
Common problems and fixes
- The parsed tree differs between machines. With Beautiful Soup, different installed backends can change the default parser. Specify the backend, keep dependencies aligned, and test the chosen behavior.
- A selector finds nothing on malformed markup. Inspect the parsed tree and compare backend behavior. Use Beautiful Soup’s
diagnose()helper when applicable, then choose parsing rules that match your intended result. - Beautiful Soup is too slow for a critical path. Its documentation says it will not be as fast as the parsers it sits on top of. Try its lxml backend first; for response-time-critical code, evaluate using lxml directly.
- The result omits content added by JavaScript. A parser does not render JavaScript. Parsing HTML source alone does not establish what a browser would show after scripts run; use a rendering or capture workflow if the rendered page is the required input.
- A benchmark winner is slower on your job. The published selectolax figures cover a specific task. Benchmark representative input and identical extraction work in your own environment before choosing on speed.
Frequently asked questions
Is Beautiful Soup itself an HTML parser?
It is a Python-facing parsing and extraction interface that delegates markup parsing to a selected backend. The backend affects the resulting tree.
Which option follows browser-style HTML parsing rules?
html5lib is designed to conform to the WHATWG HTML specification. If that behavior is essential, evaluate it with your input and chosen tree builder.
Can a Python HTML parser retrieve content created by JavaScript?
No. Parsing markup and rendering a page in a browser are separate tasks; a parser alone does not execute page scripts.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

