DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Beautiful Soup

How to Find HTML Elements by Attribute Using BeautifulSoup

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Beautiful Soup’s find() or find_all() with an attribute filter. For example, soup.find_all("a", attrs={"data-id": "42"}) returns every link whose data-id is exactly 42; find() returns only the first match. Keyword arguments such as id, type, and class_ cover common attributes, while the attrs dictionary works for hyphenated, reserved, and unusual names.

Install Beautiful Soup and parse the HTML

Install the parser and an HTML parser implementation in the environment where your script runs:

python -m pip install beautifulsoup4 lxml

Then create a BeautifulSoup object. The built-in html.parser needs no additional package; lxml is another supported parser.

from bs4 import BeautifulSoup

html = '''
Answer
Other
'''
soup = BeautifulSoup(html, "html.parser")

Beautiful Soup parses the markup into a searchable tree. Attribute searches inspect that parsed tree; they do not execute JavaScript or download content that is absent from the HTML you provide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find one element or every matching element

Use find() for the first match

find() returns the first tag that satisfies all supplied conditions, or None when nothing matches.

answer = soup.find("a", attrs={"data-id": "42"})
if answer is not None:
    print(answer.get_text(strip=True))
    print(answer.get("href"))

The result is a Beautiful Soup Tag. Read an attribute with tag.get("attribute"); this safely returns None when the attribute is missing. Indexing, such as tag["href"], raises KeyError if the attribute does not exist.

Use find_all() for all matches

find_all() returns a list-like ResultSet. An empty result means no tag met every condition.

links = soup.find_all("a", attrs={"data-id": "42"})
for link in links:
    print(link.get_text(strip=True), link.get("href"))

You can limit the number of results with limit:

first_two = soup.find_all("a", limit=2)

Filter common attributes with keyword arguments

Beautiful Soup lets you pass many attribute names directly as keyword arguments. The tag name is the first positional argument.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
main = soup.find("div", id="main")
email_fields = soup.find_all("input", type="email")
checked = soup.find_all("input", disabled=True)

These forms are concise when the attribute name is a valid Python keyword and does not conflict with a Beautiful Soup parameter. Attribute values can be exact strings, lists, regular expressions, callables, True, or None.

Search by id

content = soup.find(id="content")
# Equivalent explicit form:
content = soup.find(attrs={"id": "content"})

IDs should normally be unique in valid HTML, but find_all(id="content") is useful when dealing with imperfect documents.

Search by class with class_

Python reserves the word class, so Beautiful Soup uses class_:

cards = soup.find_all("div", class_="card")

Beautiful Soup treats HTML classes as multiple tokens. Therefore class_="body" matches <p class="body strikeout"> because one token matches. An exact string such as class_="body strikeout" is order-sensitive and requires that serialized value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For multiple classes regardless of their order, use a CSS selector:

matches = soup.select("p.body.strikeout")

The Beautiful Soup documentation notes that CSS-class searching with class_ is available as of Beautiful Soup 4.1.2.

Use attrs for any attribute name

The attrs dictionary is the reliable option for hyphenated names, ARIA attributes, data-* attributes, and names that conflict with Beautiful Soup’s arguments.

data_items = soup.find_all(attrs={"data-test-id": "checkout"})
close_buttons = soup.find_all(attrs={"aria-label": "Close"})
fields = soup.find_all(attrs={"name": "email"})

name is especially important: Beautiful Soup uses the first name argument to mean a tag name, so search an HTML name attribute through attrs={"name": ...}.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Combine a tag name and several attributes

buttons = soup.find_all(
    "button",
    attrs={"data-action": "save", "aria-label": "Save changes"}
)

All conditions are applied together. A tag must be a button and have both requested attribute values.

Match exact, alternative, present, or missing values

Exact string

product = soup.find_all("div", attrs={"data-state": "open"})

Any of several values

Pass a list when any listed value is acceptable:

active = soup.find_all(attrs={"data-state": ["open", "active"]})

Attribute presence

Use True to require that an attribute exists, regardless of its value:

disabled_controls = soup.find_all(attrs={"disabled": True})

This also handles boolean HTML attributes such as disabled, required, and checked.

Attribute absence

Use None to match tags where the attribute is not present:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
without_title = soup.find_all("a", attrs={"title": None})

Regular expressions

Regular expressions are useful for URL prefixes, identifier patterns, and other flexible matches:

import re

product_links = soup.find_all("a", href=re.compile(r"^/products/"))
data_ids = soup.find_all(attrs={"data-id": re.compile(r"^item-d+$")})

The expression is evaluated against the attribute value. Keep patterns specific enough to avoid collecting unrelated links.

Callable predicates

A function receives the candidate attribute value. Guard against None before calling string methods:

menu_items = soup.find_all(
    attrs={"aria-label": lambda value: value and "menu" in value.lower()}
)

Callables are appropriate when matching requires normalization, a numeric test, or several rules that are awkward in a regular expression.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find every data-* attribute

When the attribute name is known, use an exact attrs filter:

rows = soup.find_all(attrs={"data-row": True})
open_cards = soup.find_all(attrs={"data-state": "open"})

To discover all tags carrying any attribute whose name starts with data-, inspect each tag’s attribute mapping:

for tag in soup.find_all(True):
    data_attributes = {
        key: value for key, value in tag.attrs.items()
        if key.startswith("data-")
    }
    if data_attributes:
        print(tag.name, data_attributes)

find_all(True) visits every tag. This approach is preferable when you need the names and values of all custom data attributes rather than one known key.

Use CSS selectors for combined conditions

select() uses SoupSieve, Beautiful Soup’s CSS-selector engine. It is often clearer for multiple classes, descendants, sibling relationships, and attribute operators.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
home = soup.select('a[href="/home"]')
role_cards = soup.select('[data-role="card"]')
headlines = soup.select('article[data-kind="news"] h2 a')

Common attribute selectors include:

Selector Meaning Example
[attr] Attribute is present [disabled]
[attr="value"] Exact value [data-state="open"]
[attr^="prefix"] Value starts with text [href^="/products/"]
[attr$="suffix"] Value ends with text [href$=".pdf"]
[attr*="part"] Value contains text [aria-label*="menu"]
.one.two Both class tokens p.body.strikeout

CSS selectors return a list. Use soup.select_one(...) when you want only the first matching tag.

Choose the right method

Need Best fit Reason
One predictable result find() or select_one() Returns the first match or None.
Every matching tag find_all() or select() Returns all matches.
Simple exact attribute Keyword argument or attrs Readable and direct.
Hyphenated or reserved name attrs Avoids Python and Beautiful Soup naming conflicts.
Regex, alternatives, presence, or custom logic attrs with a value, list, True, None, regex, or callable Expresses flexible predicates.
Multiple classes or structure select() CSS syntax expresses combined conditions compactly.

Extract values safely after matching

Finding a tag and extracting its value are separate operations. Use get_text(strip=True) for visible text and get() for optional attributes.

for card in soup.select('article[data-kind="news"]'):
    title = card.select_one("h2")
    link = card.select_one("a[href]")
    print({
        "title": title.get_text(" ", strip=True) if title else None,
        "href": link.get("href") if link else None,
    })

Do not assume a selector always matches. Check for None before calling methods on a result from find() or select_one().

Complete example: filter forms and links by attributes

from bs4 import BeautifulSoup
import re

html = """
One Help """ soup = BeautifulSoup(html, "html.parser") form = soup.find("form", id="signup") email = soup.find("input", attrs={"name": "email", "type": "email"}) required = soup.find_all("input", attrs={"required": True}) product_links = soup.find_all("a", href=re.compile(r"^/products/")) submit = soup.select_one('button[aria-label*="Submit"]') print(form.get("data-state") if form else None) print(email.get("name") if email else None) print(len(required)) print([link.get_text(strip=True) for link in product_links]) print(submit.get_text(strip=True) if submit else None)
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting attribute searches

The result is None or an empty list

  • Print or save the HTML you actually parsed. The browser’s Elements panel may show DOM added later by JavaScript, while your downloaded response contains only the initial markup.
  • Check spelling, capitalization, hyphens, and the exact attribute value.
  • Verify that you supplied the correct tag name. Remove the tag constraint temporarily with soup.find_all(attrs={...}) to see whether the attribute exists elsewhere.
  • Confirm that you are searching the intended document, not an error page, login page, or redirect response.

class= raises a syntax error

Use class_="token", or use attrs={"class": "token"}. Never write class= as a Python keyword argument.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A name filter behaves unexpectedly

Use attrs={"name": "value"}. The positional or keyword name parameter is reserved for selecting tag names.

A class combination misses elements

HTML class order can vary and classes can be separated into multiple tokens. Prefer soup.select(".first.second") when both classes are required.

A callable raises an exception

Missing attributes are represented by None. Test the value before using .lower(), startswith(), or other string methods.

The page needs JavaScript

Beautiful Soup is an HTML parser, not a browser. Fetch the page with an appropriate HTTP client, use a browser automation tool when JavaScript must run, or locate an underlying data endpoint permitted by the site’s terms. Then pass the resulting HTML to Beautiful Soup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability, and responsible scraping

Keep searches narrow: specify a tag, attribute, or container instead of scanning every tag repeatedly. Parse once, reuse the soup object, and compile a regular expression once when applying it in a loop. For large documents, select a containing section first and search within that smaller tag.

Selectors are only as stable as the markup they target. Semantic attributes such as data-testid, aria-label, or documented IDs are generally less fragile than positional selectors. Treat missing matches as a normal condition, log the URL and selector, and write tests against representative HTML fixtures.

Respect a site’s terms, robots guidance, rate limits, privacy rules, and copyright. Identify your client where appropriate, avoid collecting unnecessary personal data, and cache responses when permitted.

Or skip the browser setup

If your goal is to obtain a clean visual record of a page rather than parse its source, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request can return PNG, JPEG, WebP, or PDF:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all parameters. Cookie and consent banners are accepted before capture, and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and whether it was billed. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—let Claude, Cursor, and other MCP clients capture pages.

The Free plan includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account to get started.

Frequently asked questions

Can I search an attribute without specifying a tag?

Yes. Call soup.find_all(attrs={"data-id": "42"}) or use a CSS selector such as soup.select('[data-id="42"]').

How do I match a case-insensitive attribute?

Use a callable that normalizes the value, for example lambda value: value and value.lower() == "close", while guarding against None.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Beautiful Soup download a web page for me?

No. Supply it with HTML obtained by your HTTP client, file, or browser automation workflow; it parses the content you pass in.

What does an attribute with multiple values look like?

The class attribute is represented as a list of class tokens. Other attributes are normally strings unless the parser and HTML semantics define otherwise.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.