Free tools Windows power users keep installed
One-click scans. No signup required.
Use Beautiful Soup’s find() or find_all() with an attribute filter. For example, soup.find_all("a", attrs={"data-id": "42"}) returns every link whose data-id is exactly 42; find() returns only the first match. Keyword arguments such as id, type, and class_ cover common attributes, while the attrs dictionary works for hyphenated, reserved, and unusual names.
Install Beautiful Soup and parse the HTML
Install the parser and an HTML parser implementation in the environment where your script runs:
python -m pip install beautifulsoup4 lxml
Then create a BeautifulSoup object. The built-in html.parser needs no additional package; lxml is another supported parser.
from bs4 import BeautifulSoup
html = '''
Answer
Other
'''
soup = BeautifulSoup(html, "html.parser")
Beautiful Soup parses the markup into a searchable tree. Attribute searches inspect that parsed tree; they do not execute JavaScript or download content that is absent from the HTML you provide.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Find one element or every matching element
Use find() for the first match
find() returns the first tag that satisfies all supplied conditions, or None when nothing matches.
answer = soup.find("a", attrs={"data-id": "42"})
if answer is not None:
print(answer.get_text(strip=True))
print(answer.get("href"))
The result is a Beautiful Soup Tag. Read an attribute with tag.get("attribute"); this safely returns None when the attribute is missing. Indexing, such as tag["href"], raises KeyError if the attribute does not exist.
Use find_all() for all matches
find_all() returns a list-like ResultSet. An empty result means no tag met every condition.
links = soup.find_all("a", attrs={"data-id": "42"})
for link in links:
print(link.get_text(strip=True), link.get("href"))
You can limit the number of results with limit:
first_two = soup.find_all("a", limit=2)
Filter common attributes with keyword arguments
Beautiful Soup lets you pass many attribute names directly as keyword arguments. The tag name is the first positional argument.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minutemain = soup.find("div", id="main")
email_fields = soup.find_all("input", type="email")
checked = soup.find_all("input", disabled=True)
These forms are concise when the attribute name is a valid Python keyword and does not conflict with a Beautiful Soup parameter. Attribute values can be exact strings, lists, regular expressions, callables, True, or None.
Search by id
content = soup.find(id="content")
# Equivalent explicit form:
content = soup.find(attrs={"id": "content"})
IDs should normally be unique in valid HTML, but find_all(id="content") is useful when dealing with imperfect documents.
Search by class with class_
Python reserves the word class, so Beautiful Soup uses class_:
cards = soup.find_all("div", class_="card")
Beautiful Soup treats HTML classes as multiple tokens. Therefore class_="body" matches <p class="body strikeout"> because one token matches. An exact string such as class_="body strikeout" is order-sensitive and requires that serialized value.
Rank #2
For multiple classes regardless of their order, use a CSS selector:
matches = soup.select("p.body.strikeout")
The Beautiful Soup documentation notes that CSS-class searching with class_ is available as of Beautiful Soup 4.1.2.
Use attrs for any attribute name
The attrs dictionary is the reliable option for hyphenated names, ARIA attributes, data-* attributes, and names that conflict with Beautiful Soup’s arguments.
data_items = soup.find_all(attrs={"data-test-id": "checkout"})
close_buttons = soup.find_all(attrs={"aria-label": "Close"})
fields = soup.find_all(attrs={"name": "email"})
name is especially important: Beautiful Soup uses the first name argument to mean a tag name, so search an HTML name attribute through attrs={"name": ...}.
Combine a tag name and several attributes
buttons = soup.find_all(
"button",
attrs={"data-action": "save", "aria-label": "Save changes"}
)
All conditions are applied together. A tag must be a button and have both requested attribute values.
Match exact, alternative, present, or missing values
Exact string
product = soup.find_all("div", attrs={"data-state": "open"})
Any of several values
Pass a list when any listed value is acceptable:
active = soup.find_all(attrs={"data-state": ["open", "active"]})
Attribute presence
Use True to require that an attribute exists, regardless of its value:
disabled_controls = soup.find_all(attrs={"disabled": True})
This also handles boolean HTML attributes such as disabled, required, and checked.
Attribute absence
Use None to match tags where the attribute is not present:
Recommended Free Tools
without_title = soup.find_all("a", attrs={"title": None})
Regular expressions
Regular expressions are useful for URL prefixes, identifier patterns, and other flexible matches:
import re
product_links = soup.find_all("a", href=re.compile(r"^/products/"))
data_ids = soup.find_all(attrs={"data-id": re.compile(r"^item-d+$")})
The expression is evaluated against the attribute value. Keep patterns specific enough to avoid collecting unrelated links.
Callable predicates
A function receives the candidate attribute value. Guard against None before calling string methods:
menu_items = soup.find_all(
attrs={"aria-label": lambda value: value and "menu" in value.lower()}
)
Callables are appropriate when matching requires normalization, a numeric test, or several rules that are awkward in a regular expression.
Find every data-* attribute
When the attribute name is known, use an exact attrs filter:
rows = soup.find_all(attrs={"data-row": True})
open_cards = soup.find_all(attrs={"data-state": "open"})
To discover all tags carrying any attribute whose name starts with data-, inspect each tag’s attribute mapping:
for tag in soup.find_all(True):
data_attributes = {
key: value for key, value in tag.attrs.items()
if key.startswith("data-")
}
if data_attributes:
print(tag.name, data_attributes)
find_all(True) visits every tag. This approach is preferable when you need the names and values of all custom data attributes rather than one known key.
Use CSS selectors for combined conditions
select() uses SoupSieve, Beautiful Soup’s CSS-selector engine. It is often clearer for multiple classes, descendants, sibling relationships, and attribute operators.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
home = soup.select('a[href="/home"]')
role_cards = soup.select('[data-role="card"]')
headlines = soup.select('article[data-kind="news"] h2 a')
Common attribute selectors include:
| Selector | Meaning | Example |
|---|---|---|
[attr] |
Attribute is present | [disabled] |
[attr="value"] |
Exact value | [data-state="open"] |
[attr^="prefix"] |
Value starts with text | [href^="/products/"] |
[attr$="suffix"] |
Value ends with text | [href$=".pdf"] |
[attr*="part"] |
Value contains text | [aria-label*="menu"] |
.one.two |
Both class tokens | p.body.strikeout |
CSS selectors return a list. Use soup.select_one(...) when you want only the first matching tag.
Choose the right method
| Need | Best fit | Reason |
|---|---|---|
| One predictable result | find() or select_one() |
Returns the first match or None. |
| Every matching tag | find_all() or select() |
Returns all matches. |
| Simple exact attribute | Keyword argument or attrs |
Readable and direct. |
| Hyphenated or reserved name | attrs |
Avoids Python and Beautiful Soup naming conflicts. |
| Regex, alternatives, presence, or custom logic | attrs with a value, list, True, None, regex, or callable |
Expresses flexible predicates. |
| Multiple classes or structure | select() |
CSS syntax expresses combined conditions compactly. |
Extract values safely after matching
Finding a tag and extracting its value are separate operations. Use get_text(strip=True) for visible text and get() for optional attributes.
for card in soup.select('article[data-kind="news"]'):
title = card.select_one("h2")
link = card.select_one("a[href]")
print({
"title": title.get_text(" ", strip=True) if title else None,
"href": link.get("href") if link else None,
})
Do not assume a selector always matches. Check for None before calling methods on a result from find() or select_one().
Complete example: filter forms and links by attributes
from bs4 import BeautifulSoup
import re
html = """
One
Help
"""
soup = BeautifulSoup(html, "html.parser")
form = soup.find("form", id="signup")
email = soup.find("input", attrs={"name": "email", "type": "email"})
required = soup.find_all("input", attrs={"required": True})
product_links = soup.find_all("a", href=re.compile(r"^/products/"))
submit = soup.select_one('button[aria-label*="Submit"]')
print(form.get("data-state") if form else None)
print(email.get("name") if email else None)
print(len(required))
print([link.get_text(strip=True) for link in product_links])
print(submit.get_text(strip=True) if submit else None)
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting attribute searches
The result is None or an empty list
- Print or save the HTML you actually parsed. The browser’s Elements panel may show DOM added later by JavaScript, while your downloaded response contains only the initial markup.
- Check spelling, capitalization, hyphens, and the exact attribute value.
- Verify that you supplied the correct tag name. Remove the tag constraint temporarily with
soup.find_all(attrs={...})to see whether the attribute exists elsewhere. - Confirm that you are searching the intended document, not an error page, login page, or redirect response.
class= raises a syntax error
Use class_="token", or use attrs={"class": "token"}. Never write class= as a Python keyword argument.
A name filter behaves unexpectedly
Use attrs={"name": "value"}. The positional or keyword name parameter is reserved for selecting tag names.
A class combination misses elements
HTML class order can vary and classes can be separated into multiple tokens. Prefer soup.select(".first.second") when both classes are required.
A callable raises an exception
Missing attributes are represented by None. Test the value before using .lower(), startswith(), or other string methods.
The page needs JavaScript
Beautiful Soup is an HTML parser, not a browser. Fetch the page with an appropriate HTTP client, use a browser automation tool when JavaScript must run, or locate an underlying data endpoint permitted by the site’s terms. Then pass the resulting HTML to Beautiful Soup.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchBest Value
Performance, reliability, and responsible scraping
Keep searches narrow: specify a tag, attribute, or container instead of scanning every tag repeatedly. Parse once, reuse the soup object, and compile a regular expression once when applying it in a loop. For large documents, select a containing section first and search within that smaller tag.
Selectors are only as stable as the markup they target. Semantic attributes such as data-testid, aria-label, or documented IDs are generally less fragile than positional selectors. Treat missing matches as a normal condition, log the URL and selector, and write tests against representative HTML fixtures.
Respect a site’s terms, robots guidance, rate limits, privacy rules, and copyright. Identify your client where appropriate, avoid collecting unnecessary personal data, and cache responses when permitted.
Or skip the browser setup
If your goal is to obtain a clean visual record of a page rather than parse its source, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request can return PNG, JPEG, WebP, or PDF:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for all parameters. Cookie and consent banners are accepted before capture, and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and whether it was billed. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—let Claude, Cursor, and other MCP clients capture pages.
The Free plan includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account to get started.
Frequently asked questions
Can I search an attribute without specifying a tag?
Yes. Call soup.find_all(attrs={"data-id": "42"}) or use a CSS selector such as soup.select('[data-id="42"]').
How do I match a case-insensitive attribute?
Use a callable that normalizes the value, for example lambda value: value and value.lower() == "close", while guarding against None.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Does Beautiful Soup download a web page for me?
No. Supply it with HTML obtained by your HTTP client, file, or browser automation workflow; it parses the content you pass in.
What does an attribute with multiple values look like?
The class attribute is represented as a list of class tokens. Other attributes are normally strings unless the parser and HTML semantics define otherwise.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




