PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Python’s built-in html.parser when you need dependency-free, event-driven processing. Use Beautiful Soup when you need to search and navigate a document tree; select its parser backend explicitly: lxml is very fast, while html5lib most closely follows browser-style recovery of broken HTML but is very slow. The right choice depends on whether your priority is zero dependencies, convenient extraction, speed, or tolerance of malformed markup.
This guide shows complete, reproducible examples for each approach, explains what parsing can and cannot do, and gives a troubleshooting path for real projects. Parsing begins only after HTML text or a file is available; downloading a URL and rendering JavaScript are separate concerns.
What HTML parsing does
An HTML parser turns markup text into events or a tree that your program can inspect. For example, the source <h1>News</h1><p>Read this</p> contains a heading element and a paragraph. A parser identifies those elements, their attributes, and their text so code can extract, transform, or validate content.
Python’s markup-processing modules include html.parser. The Python documentation describes an HTMLParser instance as being fed HTML data and calling handler methods for start tags, end tags, text, comments, and other markup. Beautiful Soup is a higher-level library that builds a navigable tree around a parser backend.
#1 Best Overall
Choose a parser
| Choice | Best fit | Trade-offs |
|---|---|---|
html.parser |
Small scripts, standard-library deployments, handler-based processing | No third-party install; event-oriented API; it parses invalid markup but does not validate matching tags. |
Beautiful Soup + lxml |
Tree navigation when speed matters | Requires the external lxml package and its C dependency. |
Beautiful Soup + html5lib |
Browser-like recovery of badly formed HTML | Very lenient and very slow; requires an external Python package. |
Beautiful Soup + html.parser |
Convenient tree API without an additional parser package | Recovery behavior differs from the other backends on malformed input. |
Beautiful Soup’s documentation warns that different backends can create different trees for invalid markup. Therefore, pass the backend name explicitly in production code and tests.
Parse simple HTML with Python’s standard library
Collect text from selected tags
Subclass HTMLParser and implement the callbacks you need. This example collects heading and paragraph text while preserving the order in which it appears:
from html.parser import HTMLParser
class TextExtractor(HTMLParser):
def __init__(self):
super().__init__(convert_charrefs=True)
self.capture = False
self.current_tag = None
self.parts = []
self.records = []
def handle_starttag(self, tag, attrs):
if tag in {"h1", "h2", "p"}:
self.capture = True
self.current_tag = tag
self.parts = []
def handle_data(self, data):
if self.capture:
self.parts.append(data)
def handle_endtag(self, tag):
if self.capture and tag == self.current_tag:
text = " ".join("".join(self.parts).split())
if text:
self.records.append((tag, text))
self.capture = False
self.current_tag = None
self.parts = []
html = """
Parser guide
Use a handler for a small extraction task.
Details
"""
parser = TextExtractor()
parser.feed(html)
parser.close()
print(parser.records)
# [('h1', 'Parser guide'), ('p', 'Use a handler for a small extraction task.'),
# ('h2', 'Details')]
convert_charrefs=True is the documented Python 3.10 default. Character references are converted except in contexts such as script and style. Calling close() after the final feed() lets the parser finish buffered data.
Read attributes
The attrs argument is a list of (name, value) pairs. Convert it to a dictionary when duplicate attributes are not relevant to your task:
from html.parser import HTMLParser
class LinkParser(HTMLParser):
def handle_starttag(self, tag, attrs):
if tag == "a":
attributes = dict(attrs)
href = attributes.get("href")
label = attributes.get("title", "(no title)")
print(href, label)
parser = LinkParser()
parser.feed('Docs')
parser.close()
Keep the list form if duplicate attributes matter. HTMLParser is not a strict nesting validator: it does not check that end tags match start tags, and an element implicitly closed by an outer element may not trigger the corresponding end-tag callback. Treat it as a tolerant event stream, not as an HTML conformance checker.
Handle comments and declarations
from html.parser import HTMLParser
class MarkupEvents(HTMLParser):
def handle_comment(self, data):
print("comment:", data)
def handle_decl(self, decl):
print("declaration:", decl)
def handle_startendtag(self, tag, attrs):
print("self-closing:", tag, attrs)
p = MarkupEvents()
p.feed('<!DOCTYPE html><!-- note --><br/>')
p.close()
Parse HTML as a searchable tree with Beautiful Soup
Install and choose a backend
Install Beautiful Soup with the backend you intend to run:
Rank #2
python -m pip install beautifulsoup4
# Optional alternatives:
python -m pip install lxml html5lib
Pass "html.parser", "lxml", or "html5lib" explicitly. Omitting the backend can make behavior depend on which packages happen to be installed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Extract headings, links and visible text
from bs4 import BeautifulSoup
html = """
Parser guide
Choose a backend explicitly.
Documentation
"""
soup = BeautifulSoup(html, "html.parser")
print(soup.h1.get_text(" ", strip=True))
for link in soup.select("a[href]"):
print(link.get("href"), link.get_text(" ", strip=True))
print(soup.get_text(" ", strip=True))
select() accepts CSS selectors, while methods such as find(), find_all(), and attribute access support simpler searches. get_text(" ", strip=True) inserts a separator so words from adjacent elements do not run together.
Modify or remove nodes
from bs4 import BeautifulSoup
soup = BeautifulSoup('Keep
Remove ', "html.parser")
for ad in soup.select(".ad"):
ad.decompose()
soup.main.append(soup.new_tag("p"))
soup.main.p.string = "Added by the parser"
print(soup.main.decode())
Beautiful Soup converts input to Unicode and accepts either markup text or an open file handle. Its tree model is convenient for edits, but the exact tree for malformed input is backend-dependent.
Backend behavior and malformed markup
html.parser
Use it when a standard-library dependency and direct callbacks are more important than a high-level query API. It accepts imperfect markup, but its event behavior is not a browser’s tree-building algorithm and it does not verify matching start and end tags.
lxml
Use Beautiful Soup with lxml when you want the same convenient tree interface with the speed-oriented backend described in the documentation. Plan for an external C dependency in your deployment and lock the package versions used by your tests.
html5lib
Choose html5lib when browser-like HTML5 recovery is more important than runtime. The documentation characterizes it as extremely lenient and very slow, and it requires an external Python package.
Why explicit selection matters
Malformed source can produce different parent-child relationships under each backend. A selector that succeeds with one tree can fail with another. Set the backend in every BeautifulSoup(...) call, add malformed fixtures to tests, and compare the resulting tree when upgrading dependencies.
Parse a file safely
Open files with the encoding your producer specifies. If the encoding is unknown, determine it from reliable metadata rather than silently assuming one:
from pathlib import Path
from bs4 import BeautifulSoup
source = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(source, "html.parser")
title = soup.title.get_text(strip=True) if soup.title else None
print(title)
For large documents, avoid retaining unnecessary trees or text. A streaming HTMLParser handler can emit records as events arrive; Beautiful Soup generally builds a complete in-memory tree, which is easier to query but consumes more memory.
Parsing is not downloading or JavaScript rendering
These parsers operate on HTML you already have. Getting a response from a remote URL introduces separate questions: HTTP status handling, timeouts, redirects, response encoding, authentication, and robots or access policies. The parser documentation does not define a complete network-fetching workflow, so choose and document an HTTP client independently.
Likewise, parsing the original response does not execute JavaScript. If content is inserted after page load, obtain the rendered HTML with an appropriate browser automation workflow, then pass that resulting HTML to your parser. Do not assume that a parser can reveal data absent from its input.
Or skip the browser setup
If your goal is to obtain a clean snapshot of a URL before parsing, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
cURL example (see the ScreenshotNeo API documentation):
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server so Claude, Cursor, and other MCP clients can call take_screenshot, get_page_info, and capture_pdf. Free accounts include 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Troubleshooting
“Feature X is missing”
Check which object you created. HTMLParser exposes callbacks, not Beautiful Soup’s select() or find_all(). Either implement the needed callback or parse the same input with Beautiful Soup.
Different output on another machine
A backend may differ or be selected implicitly. Install the intended backend, pass its name explicitly, and pin compatible dependency versions.
Text is duplicated or runs together
HTML often contains nested formatting elements and whitespace nodes. Use get_text(" ", strip=True) in Beautiful Soup, or accumulate and normalize data in your handle_data implementation.
Expected closing-tag code never runs
HTMLParser does not call an end-tag handler for every implicit close and does not validate nesting. Track only the state your extraction requires, or use a tree parser when parent-child relationships are essential.
Expected content is absent
Print or save the exact HTML passed to the parser. The content may be loaded by JavaScript, behind authentication, or absent from the response entirely; parsing cannot manufacture it.
Best Value
Non-ASCII characters are corrupted
Decode bytes with the producer’s declared encoding before parsing, and use an explicit encoding when reading files. Do not “fix” mojibake after the tree has already been built.
Performance, reliability and maintenance
- For small, predictable extraction jobs,
html.parserminimizes installation and startup overhead. - For many CSS-style queries or edits, Beautiful Soup reduces custom state-machine code; choose
lxmlwhen its external dependency is acceptable and speed is a priority. - Use
html5libonly when its browser-like recovery justifies very slow processing. - Keep parser selection, input encoding, and extraction rules in tests. Include valid and malformed fixtures because backend recovery can change selectors.
- Limit retained nodes and stream with callbacks when document size makes a complete tree impractical.
- Separate acquisition from parsing so HTTP failures, rendering failures, and extraction failures can be diagnosed independently.
Frequently Asked Questions
Can HTMLParser validate that HTML is correctly nested?
No. Python’s HTMLParser is designed for event handling and does not check whether end tags match start tags. Use a dedicated validator when conformance checking is required.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Which Beautiful Soup parser should I use for reproducible tests?
Choose one explicitly in every call and pin the corresponding dependency versions. The best backend depends on whether your project values standard-library availability, speed, or browser-like recovery.
Can these parsers execute JavaScript?
No. They parse supplied HTML only. Obtain rendered markup separately, then parse that result.
The Bottom Line
Start with html.parser for dependency-free callbacks; choose Beautiful Soup for tree navigation, with an explicitly selected backend. Use lxml for speed when its dependency is acceptable and html5lib for browser-like recovery when performance is secondary.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →

