Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use lxml.html or etree.HTML() for HTML that may be imperfect, and use lxml.etree’s XML parser for XML and XHTML. For an in-memory document, etree.fromstring() gives you its root element; for a file or file-like source, etree.parse() returns an element tree. Then use find() or findall() for simple navigation and xpath() for more expressive queries.
This guide walks through installation, parsing, extraction, namespaces, large XML files, output, troubleshooting, and security considerations. The linked parsing guide is the lxml project’s version 5.4 documentation; exact parser defaults and behavior can vary with the lxml and libxml2 versions installed in your environment.
Install lxml in the Python environment that runs your code
Install the package into the interpreter or virtual environment that will run your script:
python -m pip install lxml
Using python -m pip helps ensure the installer is associated with that Python interpreter. The lxml project’s installation instructions describe platform-specific options. Binary wheels are available for common environments, but installation behavior and bundled native-library versions can differ by platform. On Linux, building from source requires libxml2 and libxslt development packages.
#1 Best Overall
Confirm the import works before adding parsing code:
python -c "from lxml import etree; print(etree.LXML_VERSION)"
The version tuple identifies the lxml package. When investigating a parsing or security issue, also check which libxml2 version is in use, since parser behavior depends on the native library as well.
Choose the parser to match the input
HTML and XML have different parsing expectations. Choose deliberately rather than sending every document through the same parser.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Input or use case | Recommended entry point | What to expect |
|---|---|---|
| HTML string or bytes, including imperfect markup | etree.HTML() or lxml’s HTML parser |
Attempts to recover a usable HTML tree; recovery is not a guarantee that every damaged input is preserved exactly. |
| Well-formed XML string or bytes | etree.fromstring() with XML parsing |
Returns the root element; malformed XML normally raises a parse error. |
| XML file or file-like object | etree.parse() |
Returns an ElementTree, which wraps the document tree. |
| XHTML | XML parser | XHTML follows XML rules; applying the HTML parser can produce unexpected results. |
| Very large XML input | etree.iterparse() |
Provides events while reading incrementally; it is useful when you do not want to build and retain the whole document tree at once. |
The project’s XML and HTML parsing guide explains these parsing paths. The examples below use in-memory inputs for clarity; the same distinction between HTML and XML applies when reading from files.
Parse XML from a string, bytes, or file
Parse in-memory XML
For a short XML document already held in memory, pass a bytes or string value to etree.fromstring(). The returned element is the document root:
Rank #2
from lxml import etree
xml = b"<catalog><item id='a1'>Book</item></catalog>"
root = etree.fromstring(xml)
item = root.find("item")
if item is not None:
print(item.get("id"), item.text)
This prints a1 Book. Checking whether find() returned None avoids an attribute error if an expected child is absent. For larger inputs, a bytes value can also be obtained from a file or network response before calling fromstring(); take care not to treat untrusted XML as safe merely because it is already in memory.
Parse a path or file-like source
Use etree.parse() when you want lxml to read a path or file-like object. It returns an ElementTree; call getroot() when you need the root element:
from lxml import etree
tree = etree.parse("catalog.xml")
root = tree.getroot()
for item in root.findall("item"):
print(item.get("id"), item.text)
To parse an already opened file, pass the file object instead of its path. Parsing a path can be simpler for ordinary local files; a file-like source is useful when the bytes come from another part of your program. The lxml guide covers both forms.
Parse HTML, including imperfect markup
HTML on the web may omit closing tags or contain other markup errors. The HTML parser tries to recover a useful tree rather than requiring the input to be well-formed XML:
from lxml import etree
html = "<html><body><h1>Example</h1><p>Text"
root = etree.HTML(html)
if root is not None:
headings = root.xpath("//h1/text()")
print(headings)
The example returns the heading text as a list. Recovery is a best-effort parse, not lossless repair: the resulting tree depends on the input and libxml2 recovery behavior. If the document is XHTML, use the XML parser instead of assuming that HTML recovery will preserve its XML structure or semantics.
Find elements and text with ElementPath or XPath
Use simple helpers for direct navigation
For straightforward tree navigation, find(), findall(), and findtext() accept simple ElementPath expressions:
title = root.findtext("metadata/title")
items = root.findall("item")
first_item = root.find("item")
find() returns the first match or None; findall() returns matching elements; findtext() returns the matched element’s text, or None when there is no match. These helpers are convenient when the path is simple and the result should be a tree element or its text.
Use XPath for conditions, depth, and text selection
Use .xpath() for predicates, arbitrary-depth searches, attribute conditions, and other full XPath queries:
matching = root.xpath(".//item[@id='a1']")
texts = root.xpath("//item/text()")
ids = root.xpath("//item/@id")
The type returned depends on the XPath expression. Selecting elements returns element objects; selecting text or attributes returns strings; functions such as count() return numbers, and expressions such as boolean(...) return booleans. Do not assume every XPath result can be used as an element. See the project’s XPath and XSLT guide for XPath support and examples.
Query XML namespaces correctly
Namespace-qualified XML elements are a frequent reason for an XPath query that looks right but returns no matches. Supply a separate mapping from the prefix used in your query to the namespace URI:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsfrom lxml import etree
xml = b'''<catalog xmlns="urn:example:catalog">
<item>Book</item>
</catalog>'''
tree = etree.ElementTree(etree.fromstring(xml))
ns = {"doc": "urn:example:catalog"}
items = tree.xpath("//doc:item", namespaces=ns)
print([item.text for item in items])
XPath 1.0 has no default namespace. Even though the source document writes <item> without a visible prefix, that element belongs to urn:example:catalog. Use any convenient query prefix, such as doc, and map it to the correct URI. It does not have to match a prefix used in the source document. Use the same approach for namespaced attributes or elements in more complex queries, checking the source’s namespace URI rather than guessing from how its tags are displayed.
Process a large XML document incrementally
Building a complete tree is straightforward, but it means the tree remains in memory while you process it. For large XML input, etree.iterparse() yields parsing events incrementally while building the tree. It is a blocking event iterator; if your program needs to feed data and control when parsing progresses, the lxml guide describes XMLPullParser as the pull-parsing option.
from lxml import etree
for event, elem in etree.iterparse("records.xml", events=("end",), tag="record"):
record_id = elem.get("id")
value = elem.findtext("value")
print(record_id, value)
elem.clear()
This example handles each completed record and clears its element afterward. Clearing processed elements can reduce retained tree content, but it must fit the document structure and your data needs. If later processing depends on child content or tail text, preserve what you need before clearing; parent structure may also retain references to cleared elements. Do not apply cleanup blindly to a format where later work depends on earlier tree nodes.
Serialize a parsed tree
For a byte-string representation, use etree.tostring(). Choose the method and encoding to suit the output format and the program that will consume it:
from lxml import etree
xml_bytes = etree.tostring(root, encoding="utf-8", xml_declaration=True)
with open("output.xml", "wb") as output:
output.write(xml_bytes)
Use XML serialization for XML output; if you are producing HTML, specify the HTML method where appropriate. Serialization does not make an HTML recovery parse equivalent to a faithfully preserved original document. For file output and other options, consult the lxml parsing documentation and the API reference for your installed version.
Best Value
Or skip the browser setup
If your goal is to inspect a page as rendered in a browser, rather than parse HTML you already have, ScreenshotNeo can return a screenshot or PDF from one GET request. It is a screenshot API and MCP server, not an HTML parser; use lxml when you need to query document data in Python. The API supports image and PDF capture, and its API parameters include options for waiting for a selector, a delay, or network idle.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the request and available parameters. Cookie banners, newsletter popups, and chat widgets are removed before capture; each of those cleanup steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the shot was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
Troubleshoot common lxml parsing problems
ModuleNotFoundError: No module named 'lxml': lxml is not installed in the interpreter running the script. Runpython -m pip install lxmlwith that interpreter, then check that your editor, notebook, or service uses the same environment.- Installation fails during a source build: the platform may not have a compatible binary wheel for the selected Python and operating-system combination. On Linux, the lxml installation instructions identify libxml2 and libxslt development packages as source-build requirements. Check the installation guide for platform-specific steps rather than assuming every environment installs identically.
- XML raises a syntax or parse error: XML must be well-formed. Inspect the reported line and column, check tag nesting and encoding, and confirm the input really is XML. Do not switch to the HTML parser just to suppress an XML error if the document’s format is XHTML or another XML vocabulary.
- HTML output has missing or rearranged nodes: the HTML parser attempts recovery, but does not promise exact preservation of malformed input. Inspect the parsed tree and the source markup; for XHTML, use XML parsing.
- XPath returns an empty list: check whether elements are in a namespace. If so, map a query prefix to the document’s namespace URI and use that prefix in the XPath. Also verify whether your expression is selecting elements, text, or attributes.
- An XPath result has an unexpected type: result types depend on the expression. Element selection, text selection, attribute selection, and XPath functions can return different Python types. Check the expression and handle its actual result instead of treating every value as an element.
Review parser security settings for untrusted XML
Do not treat parser defaults as a complete security policy. The current generated lxml.etree API reference documents XMLParser defaults including no_network=True and resolve_entities='internal'. The parsing guide discusses controls for DTD loading, validation, entity resolution, network access, recovery, and huge_tree. In that API reference, huge_tree disables security restrictions to support very deep trees and long text content; it is not a routine speed or compatibility switch.
Recommended Free Tools
For XML from untrusted sources, configure only capabilities your application requires, keep lxml and its native dependencies current, and check the documentation for the versions actually deployed. The parsing guide linked here is version 5.4, while the linked XPath guide is version 4.3; generated API defaults should not be assumed to describe every past or future release. Test security-sensitive behavior against your deployed lxml/libxml2 stack.
Frequently asked questions
Can lxml parse an HTML page fetched from a URL?
Yes, but keep fetching and parsing as separate responsibilities: obtain the page content using the mechanism appropriate to your application, then pass the HTML content to an HTML parser. A parser does not by itself guarantee that content requiring browser rendering, such as JavaScript-generated markup, is present in the fetched source.
Should I use lxml or a regular expression to extract HTML?
For nested markup, a parser gives you a tree and lets you select elements by path or XPath. A regular expression can match text patterns, but it does not provide the same tree navigation model for nested or malformed HTML.
Where can I check what versions my program is using?
Inspect etree.LXML_VERSION for the Python package and etree.LIBXML_VERSION for the libxml2 version exposed by the installed build. Consult documentation that matches the deployed versions when relying on exact defaults or compatibility details.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

