What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To find hyperlinks in an HTML document, parse it with BeautifulSoup, select every <a> element, and read its href attribute:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
links = [tag.get("href") for tag in soup.find_all("a")]

This returns the raw href values, including relative paths such as /about. You can resolve those values against the page URL with Python’s urllib.parse.urljoin. The basic recipe finds anchor links; URLs stored in images, scripts, forms, metadata, or other elements require separate searches.

What “all links” means in BeautifulSoup

In HTML, a normal hyperlink is represented by an <a> (anchor) element. BeautifulSoup’s find_all("a") returns every anchor in the parsed document. Calling get("href") reads its destination safely: if an anchor has no href, the result is None instead of an exception.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That distinction matters. A page can contain anchors used only as JavaScript controls, anchors with missing attributes, fragment-only links such as #pricing, mail links, telephone links, and ordinary HTTP URLs. Decide whether you want every raw value or only navigable web addresses before filtering.

Install BeautifulSoup and choose a parser

The package is installed as beautifulsoup4:

python -m pip install beautifulsoup4

BeautifulSoup can use Python’s built-in html.parser, lxml, or html5lib. Parser choice can produce different trees when markup is malformed. Specify the parser explicitly so another machine does not silently choose a different one.

Parser Dependency When it fits
html.parser Included with Python Portable scripts and small projects
lxml Install separately BeautifulSoup’s documentation lists it first among these choices when available
html5lib Install separately When HTML5-style parsing of broken markup is the priority

If you choose an external parser, install that dependency and name it in the constructor, for example BeautifulSoup(html, "lxml").

Minimal, complete example

This script parses a string, extracts every anchor’s raw href, and keeps missing attributes visible in the output:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

html = """
<main>
  <a href="/about">About</a>
  <a href="team.html">Team</a>
  <a>Button without a destination</a>
  <a href="https://example.com/docs">Documentation</a>
</main>
"""

soup = BeautifulSoup(html, "html.parser")

for anchor in soup.find_all("a"):
    print(anchor.get("href"))

The output is:

/about
team.html
None
https://example.com/docs

Use anchor["href"] only when you have already established that the attribute exists. For general documents, get("href") avoids a KeyError.

Extract a clean list instead of printing

Keep missing values

links = [anchor.get("href") for anchor in soup.find_all("a")]
print(links)

Keeping None can be useful for auditing invalid or JavaScript-only anchors.

Discard anchors without href

links = [
    href
    for anchor in soup.find_all("a")
    if (href := anchor.get("href"))
]

This removes absent and empty values. It does not decide whether a value such as mailto:[email protected] or #contact is useful; those are still non-empty hrefs.

Remove duplicates while preserving order

seen = set()
unique_links = []

for anchor in soup.find_all("a"):
    href = anchor.get("href")
    if href and href not in seen:
        seen.add(href)
        unique_links.append(href)

Do not use a plain set if document order matters. The loop preserves the first occurrence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read link text with the destination

for anchor in soup.find_all("a"):
    href = anchor.get("href")
    label = anchor.get_text(" ", strip=True)
    print({"text": label, "href": href})

Visible text can be empty when an anchor contains only an image or is controlled by script, so treat it as descriptive data rather than a guaranteed label.

Convert relative href values to absolute URLs

HTML commonly uses relative destinations. Python’s urllib.parse.urljoin combines each value with the page URL:

from urllib.parse import urljoin

page_url = "https://example.com/docs/start.html"

absolute_links = [
    urljoin(page_url, href)
    for anchor in soup.find_all("a")
    if (href := anchor.get("href"))
]

for url in absolute_links:
    print(url)

For a base of https://example.com/docs/start.html, /about becomes https://example.com/about, while team.html resolves under /docs/. A value that is already absolute remains absolute. Scheme-relative values such as //cdn.example.net/file can supply a different host and scheme.

Validate untrusted destinations

urljoin is a URL resolver, not a security policy. Because an input href can provide its own host or scheme, validate the result before using it for crawling, requests, redirects, or access-controlled workflows. A common policy is to allow only http and https, and, when appropriate, require a hostname from an approved list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.parse import urljoin, urlparse

base = "https://example.com/docs/start.html"
allowed_hosts = {"example.com"}

safe_urls = []
for anchor in soup.find_all("a"):
    href = anchor.get("href")
    if not href:
        continue
    candidate = urljoin(base, href)
    parsed = urlparse(candidate)
    if parsed.scheme in {"http", "https"} and parsed.hostname in allowed_hosts:
        safe_urls.append(candidate)

Adjust the host policy to your application. Do not assume that every href on a page points back to the page’s own site.

Fetch first, parse second

Downloading a page and parsing its HTML are separate operations. Your parser receives only the string or bytes you pass to BeautifulSoup. If a server returns an error page, a login page, or a non-HTML response, extraction can be correct for the wrong document.

Keep the boundary explicit in your code:

from bs4 import BeautifulSoup

def extract_hrefs(html, parser="html.parser"):
    soup = BeautifulSoup(html, parser)
    return [a.get("href") for a in soup.find_all("a")]

with open("page.html", encoding="utf-8") as file:
    html = file.read()

print(extract_hrefs(html))

When you add an HTTP client, check that the response is the intended HTML before passing its body to this function. A static response may not include links inserted later by client-side JavaScript; BeautifulSoup does not execute that JavaScript.

Target a section or apply filters

Search only a document region

nav = soup.find("nav")
nav_links = [] if nav is None else [a.get("href") for a in nav.find_all("a")]

Restricting the search avoids collecting links from footers, menus, or unrelated embedded content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Filter by attributes

external_candidates = soup.find_all("a", href=True)
for anchor in external_candidates:
    print(anchor["href"])

You can also filter by CSS class, ID, or another attribute:

download_links = soup.select('a[href][data-kind="download"]')
for anchor in download_links:
    print(anchor.get("href"))

Attribute filtering still returns raw href values. Resolve and validate them separately when your output requires canonical absolute URLs.

What is not an anchor link?

The basic recipe does not discover every URL-looking string in a document. For example, an image URL is normally in img[src], a form destination in form[action], a stylesheet in link[href], and a script URL in script[src]. Search each element and attribute deliberately:

image_urls = [img.get("src") for img in soup.find_all("img") if img.get("src")]
form_targets = [form.get("action") for form in soup.find_all("form") if form.get("action")]
stylesheet_urls = [tag.get("href") for tag in soup.find_all("link") if tag.get("href")]
script_urls = [tag.get("src") for tag in soup.find_all("script") if tag.get("src")]

These are different inventories with different semantics; do not label the anchor list as every URL in the HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting empty or surprising results

No links are returned

  • Confirm that the input contains <a> elements, not just visible text or JavaScript templates.
  • Print or save the exact HTML passed to BeautifulSoup. You may have parsed an error, redirect, login, or blank response.
  • Check that you are searching the correct region; a call on a missing container returns no descendants.
  • If the site creates anchors after load with JavaScript, the initial HTML alone will not contain them. Use a browser-capable capture workflow when rendered content is required.

A KeyError occurs

Replace anchor["href"] with anchor.get("href"), or filter with href=True first.

Results differ between machines

Specify the same parser everywhere and install its dependency. Malformed markup can produce different trees under html.parser, lxml, and html5lib.

Relative URLs look wrong

Pass the real page URL, including its path, to urljoin. A document URL and a site home page can produce different results for the same relative href.

Only some links appear

Check whether you intentionally limited the search to a container, filtered attributes, removed duplicates, or discarded empty values. Also inspect non-anchor URL attributes if “links” was being used to mean every resource reference.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If you need a screenshot of a rendered page rather than an HTML href inventory, ScreenshotNeo provides a website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options such as full-page capture, CSS-selector elements, custom JavaScript, waits, device presets, PDFs, signed links, caching, and bulk capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Frequently Asked Questions

Does BeautifulSoup follow links or crawl a site?

No. It parses HTML you provide and extracts attributes. Crawling requires a separate fetching, queueing, and policy layer.

Can I extract links generated by JavaScript?

Not from an initial static HTML response if those anchors are created after load. Capture or obtain the rendered DOM first, then parse that HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why use get(‘href’) instead of [‘href’]?

get() returns None when the attribute is absent, while bracket indexing raises KeyError.

Should I store raw or absolute URLs?

Store raw values when preserving source markup matters; resolve with urljoin when consumers need navigable absolute URLs, then apply your host and scheme validation policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.