Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For ordinary XML files and strings, start with Python’s built-in xml.etree.ElementTree. Choose lxml.etree when you need full XPath, XML Schema validation, XSLT, or more parser controls; choose xmltodict when a JSON-like dictionary is the output you actually need and its mapping is an acceptable loss of XML structure.

These libraries are not interchangeable in every situation. The right choice depends on how you need to query, validate, transform, or preserve the document—and whether the input is trusted.

Choose a parser by the job

Library Install Data model and querying Best fit Main trade-off
xml.etree.ElementTree Python standard library; no additional package Element and ElementTree objects; iteration and ElementPath-style queries Configuration, simple XML files, and controlled payloads Its query support is limited compared with full XPath, and it is not aimed at advanced validation or XSLT workflows.
lxml.etree Third-party package ElementTree-compatible model with full XPath 1.0 and extensions Document-heavy workflows, XML Schema validation, transformations, and complex queries Adds a dependency and a native-library surface.
xmltodict Third-party package Nested dictionaries, lists, and scalar values; access values by keys Adapters and ETL steps that immediately need JSON-like data The mapping is not an exact XML tree and is a poor fit for fidelity-sensitive XML.

ElementTree is the sensible first choice when you need a small dependency footprint and ordinary tree traversal. Reach for lxml when an actual requirement calls for its advanced XML features. Use xmltodict when dictionary-shaped data is the goal—not merely because dictionaries look simpler at first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse a file or string with ElementTree

ElementTree is in the Python standard library. ET.parse() reads a path or file-like object and returns an ElementTree; getroot() gives you its root element. For XML already held in a string, use ET.fromstring(), which returns the root element directly.

import xml.etree.ElementTree as ET

tree = ET.parse("country_data.xml")
root = tree.getroot()

root_from_text = ET.fromstring(
    "<data><item id='1'>value</item></data>"
)
for item in root_from_text.findall("item"):
    print(item.get("id"), item.text)

Use parse() when you need the full document tree, such as when you will inspect or modify multiple branches. Use fromstring() for a string or bytes value already in memory. ElementTree also provides serialization and incremental/event APIs, so you do not need a third-party package just to perform basic XML traversal.

Find elements and read values

Use find() for a single matching child, findall() for matching children, and iter() to walk matching elements below a node. findall("item") looks for matching child elements of the current node; it is not a full-document XPath query. Check for missing nodes before dereferencing them, and account for optional text or attributes rather than assuming every document has the same shape.

for item in root.iter("item"):
    item_id = item.get("id")
    value = item.text
    print(item_id, value)

ElementTree is a good match for a known configuration format, simple feed, or controlled payload whose structure you can describe in code. If a query becomes awkward because it needs broader XPath behavior, that is a reason to consider lxml rather than trying to treat ElementPath as full XPath.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use lxml for XPath, validation, and transformations

lxml.etree offers an ElementTree-compatible API for XML and HTML, plus full XPath 1.0 with extensions, XML Schema validation, XSLT, and SAX-compatible interfaces. It is the practical choice when one of those capabilities is part of the job, rather than simply a preference for a different import name.

from lxml import etree

root = etree.fromstring(xml_bytes)
rows = root.xpath("//row[@status=$status]", status="ready")

schema_doc = etree.parse("schema.xsd")
schema = etree.XMLSchema(schema_doc)
if not schema.validate(etree.ElementTree(root)):
    print(schema.error_log)

The XPath variable in this example keeps the supplied status value separate from the XPath expression. Prefer this parameterized form to building an expression by interpolating a value: untrusted input should not become executable query syntax. The schema example parses an XSD, constructs an XML Schema validator, and reports the validator’s error log if validation fails.

Choose lxml when the document workflow benefits from these features and you can take on the third-party dependency. Make parser options explicit, particularly for external entities, network access, very large trees, and compressed input. Do not assume that a parser configuration appropriate for a trusted document is suitable for an untrusted one.

Turn XML into dictionaries with xmltodict

xmltodict.parse() accepts a string, file-like object, or generator and returns nested dictionary/list/scalar data. Attributes use an @ prefix by default, text content uses #text, and repeated elements become lists. That shape can make an API adapter or an ETL step convenient when the next stage serializes data as JSON.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import xmltodict

with open("feed.xml", "rb") as fh:
    doc = xmltodict.parse(fh, process_namespaces=True)

for entry in doc["feed"].get("entry", []):
    print(entry.get("title"))

For example, a repeated entry element is represented as a list, which is why the sample loops through it. If the XML contains only one occurrence, check how your data and parser configuration represent that case before relying on a list-shaped value. Handle optional keys explicitly as well.

By default, namespace declarations are treated as ordinary attributes. Setting process_namespaces=True enables namespace processing; choose a separator and mapping policy that downstream code can keep stable. The library also provides unparse() to convert its dictionary representation back to XML. That does not make the representation a lossless XML tree: use a full XML library such as lxml when exact fidelity matters, or when you need comments, processing instructions, mixed-content ordering, schema validation, XPath, or XSLT.

Handle namespaces deliberately

An XML element’s identity includes its namespace URI, not just the prefix shown in the source. Prefixes are document-level labels; the same URI can appear under different prefixes, and a default namespace can apply even when an element has no visible prefix.

With ElementTree or lxml, bind a prefix of your choosing to the namespace URI in a query map and use that prefix in the query. Do not compare only the visible prefix or assume an unprefixed query matches a namespaced element. Test documents that use a default namespace explicitly: an unprefixed XPath expression will not match namespaced elements. With xmltodict, decide on a namespace expansion, separator, and mapping policy that the consuming code can consistently use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Process large XML without retaining the whole tree

For files too large to comfortably keep as a complete tree, ElementTree’s iterparse() can emit events as it reads. It builds the tree incrementally, but does not automatically free it incrementally. In a record-oriented document, consume end events, finish processing a record once its descendants are available, and clear elements whose content is no longer needed.

import xml.etree.ElementTree as ET

for event, elem in ET.iterparse("large.xml", events=("end",)):
    if elem.tag == "record":
        # Extract the fields needed by the application here.
        process_record(elem)
        elem.clear()

Replace process_record with application code; the snippet illustrates where to extract data, not a built-in function. Clearing a completed element is useful when its contents are no longer needed, but design the traversal around the actual document structure and verify that clearing a child will not remove data required by later processing.

iterparse() performs blocking reads. If the application needs non-blocking behavior, use a pull parser or build an asynchronous I/O design around a bounded input stream. For very large or hostile documents, put limits on input bytes, nesting depth, parse time, decompression work, and records processed. Streaming alone is not a complete resource-control strategy.

Secure parsing of untrusted XML

Treat XML from users, external services, uploads, or other uncontrolled sources as hostile. XML features such as DTDs and entity expansion, external file or network resolution, and resource-intensive input can create security or availability risks. Choose a hardened parser configuration, keep dependencies patched, and put resource limits around parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Reject or disable DTDs and entity expansion unless there is a controlled, necessary reason to allow them.
  • Prevent external file and network resolution. In lxml, configure the XML parser deliberately, including entity and network settings; do not rely on assumptions about defaults.
  • Limit input size, nesting depth, parse time, decompression work, and the number of records processed.
  • Avoid XInclude and untrusted schema locations.
  • Keep XPath and XSLT expressions under application control. Do not execute expressions supplied by users.
  • For xmltodict, keep disable_entities=True unless there is a controlled reason to change it.
  • For untrusted input, consult XML security guidance and consider the hardened parsing options recommended by the defusedxml project.

Security is a property of the whole input path, not just the parser call: a byte limit does not constrain nesting depth, and disabling external access does not by itself bound parsing time. Set controls that match the ways your application accepts and processes documents.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common XML parsing problems

The parser rejects the document

Malformed XML, including mismatched tags or invalid document structure, cannot be parsed as a well-formed tree. Check the parser error and the input at the reported location. Confirm whether the call receives text, bytes, a path, or a file-like object as intended. Do not “fix” a parsing failure by disabling security controls on untrusted input.

A query returns no elements

Check whether the element is a child or a deeper descendant: ElementTree’s findall() path is limited, while lxml provides full XPath. Then check namespace handling. A default namespace means a visually unprefixed element is still namespaced; bind and use its URI in the query rather than searching only by local name.

A value or key is missing

XML often contains optional elements and attributes. In ElementTree, get() can return no attribute value, and an element can have no text. In xmltodict, a key may be absent or an element may appear once rather than repeatedly. Handle these shapes explicitly and test against representative documents.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parsing a large file still uses too much memory

iterparse() does not free the tree automatically. Process completed records and clear elements no longer needed; also check whether the application retains extracted records or other references. If the input is compressed, bound decompression work as well as the bytes and time allowed for parsing.

Validation fails

With lxml’s XML Schema validator, inspect schema.error_log for diagnostics. Confirm that the document is checked against the intended schema and that the schema itself is loaded from a trusted, controlled location.

Or skip the browser setup

ScreenshotNeo is for capturing website screenshots and PDFs, not for parsing XML. If the separate job is to capture a page visually, its API takes a URL in one GET request. For example, adapt the target URL below to the page you want to capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. The service can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. These are screenshot features and do not replace an XML parser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For that separate capture task, ScreenshotNeo is available at the free sign-up page.

Which parser should you use?

Start with ElementTree for ordinary XML and straightforward traversal. Choose lxml when full XPath, XML Schema validation, XSLT, or richer parser controls are requirements. Choose xmltodict when a nested dictionary is the useful output and the loss of exact XML structure is acceptable. Whichever library you use, make namespace behavior explicit, plan for memory on large files, and treat untrusted XML as hostile input.

Frequently Asked Questions

Can I use ElementTree without installing a package?

Yes. xml.etree.ElementTree is part of Python’s standard library.

Does xmltodict preserve XML exactly?

No. It maps XML to dictionaries, lists, and scalar values; use a full XML library when exact fidelity is required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does iterparse make parsing non-blocking?

No. ElementTree’s iterparse() performs blocking reads.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.