Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Pyppeteer’s page.xpath() to find the element, then pass the returned ElementHandle to page.evaluate() and call the browser DOM method getAttribute(). XPath returns a list, so check that the list is not empty before reading an item:

matches = await page.xpath("//a[@class='download']")
if not matches:
    attribute_value = None
else:
    attribute_value = await page.evaluate(
        '(element) => element.getAttribute("href")',
        matches[0],
    )

This returns the link’s href string, or None when the element exists but has no href attribute. A missing element and a missing attribute are separate cases.

What Pyppeteer returns from an XPath query

In Pyppeteer, await page.xpath(expression) evaluates an XPath expression in the page and returns a Python list of ElementHandle objects. The API reference documents an empty list when no element matches and allows those handles to be passed to page.evaluate() (API Reference — Pyppeteer 0.0.25).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It does not return an attribute value directly. The normal sequence is:

  1. Write an XPath expression.
  2. Await page.xpath().
  3. Check whether the list contains a handle.
  4. Pass a handle to page.evaluate().
  5. Call the DOM element’s getAttribute() method in JavaScript.

Complete runnable example

The following script launches Chromium, opens a page, finds a download link, reads its href, and closes the browser even if an error occurs:

import asyncio
from pyppeteer import launch


async def main():
    browser = await launch(headless=True)
    try:
        page = await browser.newPage()
        await page.goto("https://example.com", {"waitUntil": "networkidle2"})

        matches = await page.xpath("//a[@class='download']")
        if not matches:
            print("No matching element")
            return

        href = await page.evaluate(
            '(element) => element.getAttribute("href")',
            matches[0],
        )
        print("href:", href)
    finally:
        await browser.close()


asyncio.get_event_loop().run_until_complete(main())

Replace the URL, XPath, and attribute name with your own values. If the page creates the element after navigation, wait for a selector or another page-state signal before calling page.xpath().

Read one matching element

Because XPath can match several nodes, choose the item you need explicitly. The first result is index 0:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
matches = await page.xpath("//button[@data-action='download']")
value = None
if matches:
    value = await page.evaluate(
        '(element) => element.getAttribute("data-file")',
        matches[0],
    )

Use a more specific XPath when “first” is not a reliable meaning. For example, (//a[contains(@class, 'download')])[2] selects the second matching link, while //main//a[@aria-label='Download'] limits the search to the main content.

Distinguish no match from no attribute

Keep these outcomes separate in your code:

  • No match: matches is an empty list. There is no element handle to inspect.
  • Match with an absent attribute: getAttribute() returns JavaScript null, which Pyppeteer exposes as Python None.
  • Match with an empty attribute: an attribute such as href="" returns an empty string, not None.
matches = await page.xpath("//img[@class='hero']")
if not matches:
    result = {"status": "element-not-found", "value": None}
else:
    alt = await page.evaluate(
        '(element) => element.getAttribute("alt")',
        matches[0],
    )
    result = {"status": "found", "value": alt}

Get an attribute from every XPath match

For all matching elements, evaluate once per handle and preserve the page order:

matches = await page.xpath("//a[@class='download']")
values = [
    await page.evaluate(
        '(element) => element.getAttribute("href")',
        element,
    )
    for element in matches
]
print(values)

This list can contain strings, empty strings, and None. Do not filter out false-y values unless that is intentional; an empty attribute and an absent attribute convey different information.

Return several fields for each element

One browser evaluation can read multiple attributes from each handle:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
records = []
for element in matches:
    record = await page.evaluate(
        '(element) => ({n'
        '  href: element.getAttribute("href"),n'
        '  text: element.textContent.trim(),n'
        '  download: element.getAttribute("download")n'
        '})',
        element,
    )
    records.append(record)

Passing the handle as an argument is important: the JavaScript runs in the browser context, while the Python loop runs in your program.

Pyppeteer names versus JavaScript Puppeteer

JavaScript Puppeteer examples commonly use page.$x(). Python cannot use a dollar sign in a method name, so Pyppeteer provides page.xpath() and the shorthand page.Jx(). The project documentation describes this mapping (Pyppeteer’s documentation; project README).

matches = await page.Jx("//a[@rel='next']")

page.xpath() is usually clearer in shared code because its purpose is immediately visible. The returned value and empty-list behavior are the same.

Writing the evaluate callback correctly

The callback in the main example is a JavaScript arrow function:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
'(element) => element.getAttribute("href")'

Pyppeteer detects whether a string passed to evaluate() is a function or an expression. The documentation notes that an expression can be forced with force_expr=True if it is misdetected (Pyppeteer’s documentation).

Use force_expr for a JavaScript expression

If you deliberately pass an expression rather than a function, make that explicit:

value = await page.evaluate(
    'element => element.getAttribute("data-id")',
    matches[0],
    force_expr=False,
)

The arrow-function form is already intended to be recognized as a function, so the primary example does not need force_expr. Use the option only when the string you supply is an expression whose detection is ambiguous in your installed version.

Useful XPath patterns for attributes

Match an exact attribute

matches = await page.xpath("//input[@name='email']")

Match an attribute that contains text

matches = await page.xpath("//a[contains(@href, '/download/')]")

Match an attribute regardless of surrounding whitespace

matches = await page.xpath(
    "//button[normalize-space(@aria-label)='Close']"
)

Select an element by visible text and read another attribute

matches = await page.xpath("//a[normalize-space()='Documentation']")
if matches:
    url = await page.evaluate(
        '(element) => element.getAttribute("href")',
        matches[0],
    )

XPath predicates decide which handles you receive; getAttribute() decides which value you extract after the match.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Timing, stale handles, and dynamic pages

An XPath query only sees the DOM at the moment it runs. Query after the page has rendered the target element, and avoid retaining a handle across an action that replaces that element.

  • Navigate with an appropriate wait condition.
  • Wait for a known selector when the target is rendered asynchronously.
  • After clicking a control that redraws a list, run page.xpath() again instead of reusing old handles.
  • If a single-page application changes the attribute, evaluate after the change has completed.

A handle may become unusable when its underlying node is detached. Re-querying is safer than assuming a previously obtained handle still points to the current DOM.

Troubleshooting

“list index out of range”

Cause: the XPath returned no elements and code indexed matches[0]. Fix: test if matches: first, then handle the not-found case.

The value is always None

Cause: the element matches, but it does not have the requested attribute, or the attribute name is wrong. Inspect the element in browser developer tools and verify the exact spelling. Remember that HTML attribute names are generally case-insensitive, while the value itself is not changed by getAttribute().

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The XPath works in the inspector but not in Pyppeteer

Cause: the page or frame differs, the element has not rendered, or the expression is being evaluated against the wrong document. Wait for rendering and, when the target is inside an iframe, obtain the correct frame before querying it. Also check that Python string quoting has not altered the XPath.

JavaScript evaluation reports an invalid argument

Cause: the object passed to evaluate() is not the matched ElementHandle, or the handle has been detached. Pass the handle directly as the second argument and query again after DOM replacement.

The callback is treated as an expression

Cause: Pyppeteer’s function-versus-expression detection could not classify the supplied string. Use a clear function string such as '(element) => element.getAttribute("href")'; for an actual expression, use the documented force_expr=True option.

The script hangs during launch or navigation

Cause: Chromium startup, network activity, or a page that never reaches the selected wait condition. Set a practical navigation timeout, choose a wait condition appropriate to the site, and always close the browser in a finally block. A timeout should be treated as a failed capture, not as evidence that the attribute is absent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance and reliability choices

For a small number of matches, one evaluation per handle is straightforward and easy to debug. If a page contains thousands of matching nodes, repeated Python-to-browser round trips can become the dominant cost. In that case, you can evaluate a page-side query only after verifying the behavior against your installed Pyppeteer version; the documented guarantee used here is that an individual ElementHandle can be passed to evaluate(), not that a list of handles will be serialized as arguments.

Stable locators improve reliability. Prefer a semantic attribute, a scoped container, or a role-like marker over brittle absolute paths such as /html/body/div[3]/div[2]. Treat an empty result as a state your program handles explicitly, and log the XPath and page URL so failures can be reproduced without exposing secrets from cookies or authorization headers.

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than DOM-level extraction, ScreenshotNeo provides a website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be disabled individually. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for parameters and response details. The same request in Python is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

In Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account.

Quick decision guide

  • Need the actual DOM attribute, conditional logic, or values from many nodes? Use page.xpath() plus page.evaluate().
  • Need a screenshot or PDF without managing Chromium, consent banners, or overlays? Use ScreenshotNeo’s API.
  • Need an AI agent to perform captures? Use ScreenshotNeo’s MCP tools.

Frequently Asked Questions

Can Pyppeteer XPath return an attribute directly?

No. It returns a list of ElementHandle objects. Read the attribute by passing a selected handle to page.evaluate() and calling getAttribute().

What does getAttribute() return when the attribute is missing?

For a matched element, the browser DOM method returns null, exposed to Python as None. That differs from an empty string, which represents an attribute that exists with no value.

Should I use page.xpath() or page.Jx()?

They are Pyppeteer names for XPath lookup. page.xpath() is generally clearer; page.Jx() is the shorthand documented by the project.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.