Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Pyppeteer’s page.xpath() to find the element, then pass the returned ElementHandle to page.evaluate() and call the browser DOM method getAttribute(). XPath returns a list, so check that the list is not empty before reading an item:
matches = await page.xpath("//a[@class='download']")
if not matches:
attribute_value = None
else:
attribute_value = await page.evaluate(
'(element) => element.getAttribute("href")',
matches[0],
)
This returns the link’s href string, or None when the element exists but has no href attribute. A missing element and a missing attribute are separate cases.
What Pyppeteer returns from an XPath query
In Pyppeteer, await page.xpath(expression) evaluates an XPath expression in the page and returns a Python list of ElementHandle objects. The API reference documents an empty list when no element matches and allows those handles to be passed to page.evaluate() (API Reference — Pyppeteer 0.0.25).
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →It does not return an attribute value directly. The normal sequence is:
#1 Best Overall
- Write an XPath expression.
- Await
page.xpath(). - Check whether the list contains a handle.
- Pass a handle to
page.evaluate(). - Call the DOM element’s
getAttribute()method in JavaScript.
Complete runnable example
The following script launches Chromium, opens a page, finds a download link, reads its href, and closes the browser even if an error occurs:
import asyncio
from pyppeteer import launch
async def main():
browser = await launch(headless=True)
try:
page = await browser.newPage()
await page.goto("https://example.com", {"waitUntil": "networkidle2"})
matches = await page.xpath("//a[@class='download']")
if not matches:
print("No matching element")
return
href = await page.evaluate(
'(element) => element.getAttribute("href")',
matches[0],
)
print("href:", href)
finally:
await browser.close()
asyncio.get_event_loop().run_until_complete(main())
Replace the URL, XPath, and attribute name with your own values. If the page creates the element after navigation, wait for a selector or another page-state signal before calling page.xpath().
Read one matching element
Because XPath can match several nodes, choose the item you need explicitly. The first result is index 0:
matches = await page.xpath("//button[@data-action='download']")
value = None
if matches:
value = await page.evaluate(
'(element) => element.getAttribute("data-file")',
matches[0],
)
Use a more specific XPath when “first” is not a reliable meaning. For example, (//a[contains(@class, 'download')])[2] selects the second matching link, while //main//a[@aria-label='Download'] limits the search to the main content.
Distinguish no match from no attribute
Keep these outcomes separate in your code:
- No match:
matchesis an empty list. There is no element handle to inspect. - Match with an absent attribute:
getAttribute()returns JavaScriptnull, which Pyppeteer exposes as PythonNone. - Match with an empty attribute: an attribute such as
href=""returns an empty string, notNone.
matches = await page.xpath("//img[@class='hero']")
if not matches:
result = {"status": "element-not-found", "value": None}
else:
alt = await page.evaluate(
'(element) => element.getAttribute("alt")',
matches[0],
)
result = {"status": "found", "value": alt}
Get an attribute from every XPath match
For all matching elements, evaluate once per handle and preserve the page order:
matches = await page.xpath("//a[@class='download']")
values = [
await page.evaluate(
'(element) => element.getAttribute("href")',
element,
)
for element in matches
]
print(values)
This list can contain strings, empty strings, and None. Do not filter out false-y values unless that is intentional; an empty attribute and an absent attribute convey different information.
Rank #2
Return several fields for each element
One browser evaluation can read multiple attributes from each handle:
records = []
for element in matches:
record = await page.evaluate(
'(element) => ({n'
' href: element.getAttribute("href"),n'
' text: element.textContent.trim(),n'
' download: element.getAttribute("download")n'
'})',
element,
)
records.append(record)
Passing the handle as an argument is important: the JavaScript runs in the browser context, while the Python loop runs in your program.
Pyppeteer names versus JavaScript Puppeteer
JavaScript Puppeteer examples commonly use page.$x(). Python cannot use a dollar sign in a method name, so Pyppeteer provides page.xpath() and the shorthand page.Jx(). The project documentation describes this mapping (Pyppeteer’s documentation; project README).
matches = await page.Jx("//a[@rel='next']")
page.xpath() is usually clearer in shared code because its purpose is immediately visible. The returned value and empty-list behavior are the same.
Writing the evaluate callback correctly
The callback in the main example is a JavaScript arrow function:
Free tools Windows power users keep installed
One-click scans. No signup required.
'(element) => element.getAttribute("href")'
Pyppeteer detects whether a string passed to evaluate() is a function or an expression. The documentation notes that an expression can be forced with force_expr=True if it is misdetected (Pyppeteer’s documentation).
Use force_expr for a JavaScript expression
If you deliberately pass an expression rather than a function, make that explicit:
value = await page.evaluate(
'element => element.getAttribute("data-id")',
matches[0],
force_expr=False,
)
The arrow-function form is already intended to be recognized as a function, so the primary example does not need force_expr. Use the option only when the string you supply is an expression whose detection is ambiguous in your installed version.
Useful XPath patterns for attributes
Match an exact attribute
matches = await page.xpath("//input[@name='email']")
Match an attribute that contains text
matches = await page.xpath("//a[contains(@href, '/download/')]")
Match an attribute regardless of surrounding whitespace
matches = await page.xpath(
"//button[normalize-space(@aria-label)='Close']"
)
Select an element by visible text and read another attribute
matches = await page.xpath("//a[normalize-space()='Documentation']")
if matches:
url = await page.evaluate(
'(element) => element.getAttribute("href")',
matches[0],
)
XPath predicates decide which handles you receive; getAttribute() decides which value you extract after the match.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Timing, stale handles, and dynamic pages
An XPath query only sees the DOM at the moment it runs. Query after the page has rendered the target element, and avoid retaining a handle across an action that replaces that element.
- Navigate with an appropriate wait condition.
- Wait for a known selector when the target is rendered asynchronously.
- After clicking a control that redraws a list, run
page.xpath()again instead of reusing old handles. - If a single-page application changes the attribute, evaluate after the change has completed.
A handle may become unusable when its underlying node is detached. Re-querying is safer than assuming a previously obtained handle still points to the current DOM.
Troubleshooting
“list index out of range”
Cause: the XPath returned no elements and code indexed matches[0]. Fix: test if matches: first, then handle the not-found case.
The value is always None
Cause: the element matches, but it does not have the requested attribute, or the attribute name is wrong. Inspect the element in browser developer tools and verify the exact spelling. Remember that HTML attribute names are generally case-insensitive, while the value itself is not changed by getAttribute().
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteThe XPath works in the inspector but not in Pyppeteer
Cause: the page or frame differs, the element has not rendered, or the expression is being evaluated against the wrong document. Wait for rendering and, when the target is inside an iframe, obtain the correct frame before querying it. Also check that Python string quoting has not altered the XPath.
JavaScript evaluation reports an invalid argument
Cause: the object passed to evaluate() is not the matched ElementHandle, or the handle has been detached. Pass the handle directly as the second argument and query again after DOM replacement.
The callback is treated as an expression
Cause: Pyppeteer’s function-versus-expression detection could not classify the supplied string. Use a clear function string such as '(element) => element.getAttribute("href")'; for an actual expression, use the documented force_expr=True option.
The script hangs during launch or navigation
Cause: Chromium startup, network activity, or a page that never reaches the selected wait condition. Set a practical navigation timeout, choose a wait condition appropriate to the site, and always close the browser in a finally block. A timeout should be treated as a failed capture, not as evidence that the attribute is absent.
Performance and reliability choices
For a small number of matches, one evaluation per handle is straightforward and easy to debug. If a page contains thousands of matching nodes, repeated Python-to-browser round trips can become the dominant cost. In that case, you can evaluate a page-side query only after verifying the behavior against your installed Pyppeteer version; the documented guarantee used here is that an individual ElementHandle can be passed to evaluate(), not that a list of handles will be serialized as arguments.
Best Value
Stable locators improve reliability. Prefer a semantic attribute, a scoped container, or a role-like marker over brittle absolute paths such as /html/body/div[3]/div[2]. Treat an empty result as a state your program handles explicitly, and log the XPath and page URL so failures can be reproduced without exposing secrets from cookies or authorization headers.
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than DOM-level extraction, ScreenshotNeo provides a website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be disabled individually. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for parameters and response details. The same request in Python is:
Recommended Free Tools
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
In Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account.
Quick decision guide
- Need the actual DOM attribute, conditional logic, or values from many nodes? Use
page.xpath()pluspage.evaluate(). - Need a screenshot or PDF without managing Chromium, consent banners, or overlays? Use ScreenshotNeo’s API.
- Need an AI agent to perform captures? Use ScreenshotNeo’s MCP tools.
Frequently Asked Questions
Can Pyppeteer XPath return an attribute directly?
No. It returns a list of ElementHandle objects. Read the attribute by passing a selected handle to page.evaluate() and calling getAttribute().
What does getAttribute() return when the attribute is missing?
For a matched element, the browser DOM method returns null, exposed to Python as None. That differs from an empty string, which represents an attribute that exists with no value.
Should I use page.xpath() or page.Jx()?
They are Pyppeteer names for XPath lookup. page.xpath() is generally clearer; page.Jx() is the shorthand documented by the project.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

