What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To find hyperlinks in an HTML document, parse it with BeautifulSoup, select every <a> element, and read its href attribute:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
links = [tag.get("href") for tag in soup.find_all("a")]
This returns the raw href values, including relative paths such as /about. You can resolve those values against the page URL with Python’s urllib.parse.urljoin. The basic recipe finds anchor links; URLs stored in images, scripts, forms, metadata, or other elements require separate searches.
What “all links” means in BeautifulSoup
In HTML, a normal hyperlink is represented by an <a> (anchor) element. BeautifulSoup’s find_all("a") returns every anchor in the parsed document. Calling get("href") reads its destination safely: if an anchor has no href, the result is None instead of an exception.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteThat distinction matters. A page can contain anchors used only as JavaScript controls, anchors with missing attributes, fragment-only links such as #pricing, mail links, telephone links, and ordinary HTTP URLs. Decide whether you want every raw value or only navigable web addresses before filtering.
#1 Best Overall
Install BeautifulSoup and choose a parser
The package is installed as beautifulsoup4:
python -m pip install beautifulsoup4
BeautifulSoup can use Python’s built-in html.parser, lxml, or html5lib. Parser choice can produce different trees when markup is malformed. Specify the parser explicitly so another machine does not silently choose a different one.
| Parser | Dependency | When it fits |
|---|---|---|
html.parser |
Included with Python | Portable scripts and small projects |
lxml |
Install separately | BeautifulSoup’s documentation lists it first among these choices when available |
html5lib |
Install separately | When HTML5-style parsing of broken markup is the priority |
If you choose an external parser, install that dependency and name it in the constructor, for example BeautifulSoup(html, "lxml").
Minimal, complete example
This script parses a string, extracts every anchor’s raw href, and keeps missing attributes visible in the output:
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →from bs4 import BeautifulSoup
html = """
<main>
<a href="/about">About</a>
<a href="team.html">Team</a>
<a>Button without a destination</a>
<a href="https://example.com/docs">Documentation</a>
</main>
"""
soup = BeautifulSoup(html, "html.parser")
for anchor in soup.find_all("a"):
print(anchor.get("href"))
The output is:
/about
team.html
None
https://example.com/docs
Use anchor["href"] only when you have already established that the attribute exists. For general documents, get("href") avoids a KeyError.
Extract a clean list instead of printing
Keep missing values
links = [anchor.get("href") for anchor in soup.find_all("a")]
print(links)
Keeping None can be useful for auditing invalid or JavaScript-only anchors.
Rank #2
Discard anchors without href
links = [
href
for anchor in soup.find_all("a")
if (href := anchor.get("href"))
]
This removes absent and empty values. It does not decide whether a value such as mailto:[email protected] or #contact is useful; those are still non-empty hrefs.
Remove duplicates while preserving order
seen = set()
unique_links = []
for anchor in soup.find_all("a"):
href = anchor.get("href")
if href and href not in seen:
seen.add(href)
unique_links.append(href)
Do not use a plain set if document order matters. The loop preserves the first occurrence.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Read link text with the destination
for anchor in soup.find_all("a"):
href = anchor.get("href")
label = anchor.get_text(" ", strip=True)
print({"text": label, "href": href})
Visible text can be empty when an anchor contains only an image or is controlled by script, so treat it as descriptive data rather than a guaranteed label.
Convert relative href values to absolute URLs
HTML commonly uses relative destinations. Python’s urllib.parse.urljoin combines each value with the page URL:
from urllib.parse import urljoin
page_url = "https://example.com/docs/start.html"
absolute_links = [
urljoin(page_url, href)
for anchor in soup.find_all("a")
if (href := anchor.get("href"))
]
for url in absolute_links:
print(url)
For a base of https://example.com/docs/start.html, /about becomes https://example.com/about, while team.html resolves under /docs/. A value that is already absolute remains absolute. Scheme-relative values such as //cdn.example.net/file can supply a different host and scheme.
Validate untrusted destinations
urljoin is a URL resolver, not a security policy. Because an input href can provide its own host or scheme, validate the result before using it for crawling, requests, redirects, or access-controlled workflows. A common policy is to allow only http and https, and, when appropriate, require a hostname from an approved list.
from urllib.parse import urljoin, urlparse
base = "https://example.com/docs/start.html"
allowed_hosts = {"example.com"}
safe_urls = []
for anchor in soup.find_all("a"):
href = anchor.get("href")
if not href:
continue
candidate = urljoin(base, href)
parsed = urlparse(candidate)
if parsed.scheme in {"http", "https"} and parsed.hostname in allowed_hosts:
safe_urls.append(candidate)
Adjust the host policy to your application. Do not assume that every href on a page points back to the page’s own site.
Fetch first, parse second
Downloading a page and parsing its HTML are separate operations. Your parser receives only the string or bytes you pass to BeautifulSoup. If a server returns an error page, a login page, or a non-HTML response, extraction can be correct for the wrong document.
Keep the boundary explicit in your code:
from bs4 import BeautifulSoup
def extract_hrefs(html, parser="html.parser"):
soup = BeautifulSoup(html, parser)
return [a.get("href") for a in soup.find_all("a")]
with open("page.html", encoding="utf-8") as file:
html = file.read()
print(extract_hrefs(html))
When you add an HTTP client, check that the response is the intended HTML before passing its body to this function. A static response may not include links inserted later by client-side JavaScript; BeautifulSoup does not execute that JavaScript.
Target a section or apply filters
Search only a document region
nav = soup.find("nav")
nav_links = [] if nav is None else [a.get("href") for a in nav.find_all("a")]
Restricting the search avoids collecting links from footers, menus, or unrelated embedded content.
Filter by attributes
external_candidates = soup.find_all("a", href=True)
for anchor in external_candidates:
print(anchor["href"])
You can also filter by CSS class, ID, or another attribute:
download_links = soup.select('a[href][data-kind="download"]')
for anchor in download_links:
print(anchor.get("href"))
Attribute filtering still returns raw href values. Resolve and validate them separately when your output requires canonical absolute URLs.
What is not an anchor link?
The basic recipe does not discover every URL-looking string in a document. For example, an image URL is normally in img[src], a form destination in form[action], a stylesheet in link[href], and a script URL in script[src]. Search each element and attribute deliberately:
image_urls = [img.get("src") for img in soup.find_all("img") if img.get("src")]
form_targets = [form.get("action") for form in soup.find_all("form") if form.get("action")]
stylesheet_urls = [tag.get("href") for tag in soup.find_all("link") if tag.get("href")]
script_urls = [tag.get("src") for tag in soup.find_all("script") if tag.get("src")]
These are different inventories with different semantics; do not label the anchor list as every URL in the HTML.
Troubleshooting empty or surprising results
No links are returned
- Confirm that the input contains
<a>elements, not just visible text or JavaScript templates. - Print or save the exact HTML passed to BeautifulSoup. You may have parsed an error, redirect, login, or blank response.
- Check that you are searching the correct region; a call on a missing container returns no descendants.
- If the site creates anchors after load with JavaScript, the initial HTML alone will not contain them. Use a browser-capable capture workflow when rendered content is required.
A KeyError occurs
Replace anchor["href"] with anchor.get("href"), or filter with href=True first.
Best Value
Results differ between machines
Specify the same parser everywhere and install its dependency. Malformed markup can produce different trees under html.parser, lxml, and html5lib.
Relative URLs look wrong
Pass the real page URL, including its path, to urljoin. A document URL and a site home page can produce different results for the same relative href.
Only some links appear
Check whether you intentionally limited the search to a container, filtered attributes, removed duplicates, or discarded empty values. Also inspect non-anchor URL attributes if “links” was being used to mean every resource reference.
Free tools Windows power users keep installed
One-click scans. No signup required.
Or skip the browser setup
If you need a screenshot of a rendered page rather than an HTML href inventory, ScreenshotNeo provides a website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for options such as full-page capture, CSS-selector elements, custom JavaScript, waits, device presets, PDFs, signed links, caching, and bulk capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently Asked Questions
Does BeautifulSoup follow links or crawl a site?
No. It parses HTML you provide and extracts attributes. Crawling requires a separate fetching, queueing, and policy layer.
Can I extract links generated by JavaScript?
Not from an initial static HTML response if those anchors are created after load. Capture or obtain the rendered DOM first, then parse that HTML.
Why use get(‘href’) instead of [‘href’]?
get() returns None when the attribute is absent, while bracket indexing raises KeyError.
Should I store raw or absolute URLs?
Store raw values when preserving source markup matters; resolve with urljoin when consumers need navigable absolute URLs, then apply your host and scheme validation policy.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

