Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To extract links from a website, parse its HTML and read each anchor element’s href attribute. Keep the original href for accuracy, resolve relative values against the page URL when you need usable destinations, then decide how to handle fragments, non-web schemes, duplicates, and links that are only interface controls. For one page, Beautiful Soup is a straightforward Python choice; for crawling with domain and pattern filters, Scrapy’s LinkExtractor provides more control.
What a website link extractor actually extracts
Most ordinary page links are represented by an HTML <a href="…"> element. The href is the link’s destination value, but it is not necessarily a complete web address. It can be a relative path, a fragment pointing within the same document, or a non-HTTP scheme such as mailto:, tel:, sms:, or javascript:. The MDN reference for the anchor element describes these as valid href uses.
So “extract all links” can mean several different things: collect every href exactly as written; turn page-relative hrefs into absolute URLs; keep only HTTP(S) destinations; or build a crawl list after filtering and deduplication. Choose the output policy before discarding information. If provenance or auditing matters, retain the raw href alongside any normalized destination.
Extract links from one page with Beautiful Soup
Beautiful Soup’s documented basic pattern is to iterate over all anchor tags and print their href values. This example fetches one page, parses its received HTML, preserves raw href values, resolves them against the page URL, and stores fragments separately.
#1 Best Overall
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urldefrag
page_url = "https://example.com/articles/"
response = requests.get(page_url, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
results = []
for tag in soup.find_all("a", href=True):
raw_href = tag["href"].strip()
if not raw_href:
continue
absolute = urljoin(page_url, raw_href)
url_without_fragment, fragment = urldefrag(absolute)
results.append({
"raw_href": raw_href,
"url": url_without_fragment,
"fragment": fragment,
"text": tag.get_text(" ", strip=True),
})
for item in results:
print(item)
Install the dependencies first with python -m pip install requests beautifulsoup4. The equivalent core parser loop, when you already have a Beautiful Soup object named soup, is:
for link in soup.find_all('a'):
print(link.get('href'))
The more defensive version uses href=True to skip anchors with no href at all, strips surrounding whitespace, and ignores blank values. It still preserves the raw href, which can be useful when you need to inspect what the page supplied rather than only the resolved result.
Raw href versus resolved URL
A raw href such as ../pricing is meaningful only in relation to the document where it appeared. urljoin(page_url, raw_href) resolves it to a destination using the page’s base address. Keep both values if a later reviewer must know whether a URL came from an absolute link, a root-relative path, or a relative path.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →HTML can also provide a <base href="…"> element that changes how relative references resolve. If your extraction must match browser navigation exactly, inspect the document’s base element and use its applicable base URL rather than blindly assuming the fetched URL is the effective base. The example uses page_url as the base for clarity.
Fragments and query strings
A fragment follows # and identifies a location within a document, such as #installation. urldefrag in the example separates that fragment from the URL. Keep it when the target section matters; otherwise store it separately or remove it consistently. Query strings, by contrast, may carry search terms, pagination, or other state. Do not strip them indiscriminately. Remove tracking parameters only under an explicit, documented rule.
Extract links across pages with Scrapy
For a crawl rather than a single-page parse, Scrapy’s LxmlLinkExtractor is designed to extract links from responses. Its defaults cover a and area tags and their href attributes. Filters can constrain domains, patterns, extensions, and selected areas of a response. The extracted Link object includes the URL, anchor text, fragment, and nofollow information.
from scrapy.linkextractors import LinkExtractor
extractor = LinkExtractor(
allow_domains={"example.com"},
deny_extensions={"pdf", "zip"},
unique=True,
)
links = extractor.extract_links(response)
for link in links:
yield {
"url": link.url,
"text": link.text,
"fragment": link.fragment,
"nofollow": link.nofollow,
}
This is a callback-level extraction snippet: response must be a Scrapy response, and the yield belongs inside a spider callback. allow_domains limits results to the named domain; deny_extensions excludes matching file extensions; and unique=True filters duplicate links in the extractor’s result. Scrapy also supports allow/deny regular expressions, tag and attribute selection, XPath or CSS restrictions, whitespace handling, custom value processing, canonicalization, and unique filtering.
Recommended Free Tools
Canonicalization and crawl identity
Canonicalization may change the URL that would be seen by a server. Scrapy documents it as useful for duplicate checking, but it is a policy choice, not a harmless formatting step. If exact markup values or server behavior matter, retain the raw or non-canonical URL as well as any crawl identity key you create. A crawler can deduplicate on a normalized key without erasing the original link record.
Rank #3
Choose a link policy before deduplicating
There is no single correct definition of “duplicate.” The right choice depends on whether you are auditing markup, building a navigation graph, or scheduling requests.
| Decision | Options | Practical consequence |
|---|---|---|
| Scope | One received document or a multi-page crawl | Beautiful Soup handles page-level parsing; a Scrapy link extractor can filter links from crawl responses. |
| URL fidelity | Raw href, resolved absolute URL, fragment retained or separated | Keeping raw and resolved forms preserves both provenance and a usable destination. |
| Destination types | All schemes or only HTTP(S) | Filtering to web URLs removes email, phone, script, and other non-page destinations that may still be meaningful links. |
| Duplicates | Keep every occurrence, deduplicate exact strings, or normalize for crawl identity | Deduplication can hide repeated placement; preserve counts or source elements when auditing matters. |
| Metadata | Destination alone or destination plus text, fragment, and rel information | Text and attributes help explain a link’s purpose and identify nofollow links. |
| Dynamic content | Parse received HTML or render in a browser first | Beautiful Soup and Scrapy operate on the HTML they receive; their cited documentation does not promise browser rendering of JavaScript-created links. |
If your output is a navigation graph, a common policy is to resolve hrefs, retain HTTP(S) links, remove fragments from the request identity while storing them separately, and deduplicate after normalization. If it is an audit, retain each occurrence, its raw href, anchor text, and source page, then calculate normalized groups without deleting records.
Filter fake navigation and non-page href values
Some href values are not useful crawl destinations. A bare # or javascript:void(0) is often being used as a UI control rather than a genuine destination. MDN warns that bogus href values can behave unexpectedly when links are copied, dragged, opened in a new window, bookmarked, or when JavaScript fails or is disabled; for actions that are not navigation, MDN recommends a <button> instead. See MDN’s anchor element guidance.
Do not discard every unusual scheme automatically. mailto:, tel:, and sms: may be legitimate contact links. A document or download may be a relevant destination even if it is not an HTML page. Classify values according to your task, and make exclusions visible in the output or configuration.
Extracting links from JavaScript-rendered pages
A parser can only extract links from the HTML it receives. If the server-delivered markup does not contain a link and client-side JavaScript creates it later, the Beautiful Soup and Scrapy patterns above do not establish that the link exists. Check the received HTML first. If the link appears only after scripts run, use a browser-rendering step to obtain the rendered DOM, then parse that DOM or collect links through browser automation. Do not assume a static response is equivalent to the page a visitor sees.
Troubleshooting common extraction problems
- No results: Confirm the fetched response is the expected page, inspect its HTML for
<a href>elements, and check whether the links are rendered only after JavaScript runs. - Relative URLs appear unusable: Resolve them with
urljoinand the effective document base, including any applicable<base>element. - Links point to unexpected pages: Preserve and inspect the raw href, check the base URL, and verify whether the href is protocol-relative, root-relative, or path-relative.
- Blank or missing destinations appear: Exclude missing attributes with
href=Trueand skip values that become empty after trimming. - Email or phone links are mixed into page URLs: Parse the scheme and apply an explicit HTTP(S)-only rule if you are assembling a web crawl queue.
- Repeated URLs disappear: Check whether extractor uniqueness, canonicalization, or a later set conversion is removing occurrences; retain occurrence records if repetition is important.
- Links with fragments collapse together: Store fragment separately before deduplication, and choose whether fragments are significant for your use case.
- Query parameters vanish or multiply records: Avoid generic query stripping. Decide which parameters carry state and which, if any, are tracking-only before normalization.
- Files are missing from crawl results: Review extension filters such as
deny_extensions; a PDF or archive may be a valid target even if it is not a page.
Or skip the browser setup
If your task needs screenshots of pages rather than a link list, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns a PNG, JPEG, WebP, or PDF. It does not extract href values; use the parsing methods above for that job. For visual review of a page, it can remove cookie/consent banners, newsletter popups, and chat widgets before capture, with each cleanup step configurable. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP tools let AI agents take screenshots, inspect page info, and capture PDFs.
The following cURL request captures a page as WebP; replace the example URL with the page you want to inspect. See the ScreenshotNeo documentation for request options and response details.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo’s free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots. Sign up for free screenshots.
Cost and reliability considerations for extraction
Beautiful Soup parsing itself is local work, but retrieving pages introduces network latency, server errors, redirects, and timeouts. Set a request timeout, check HTTP status, and make crawl behavior respectful of the target site’s rules and capacity. For multi-page work, a crawler gives you scope and filtering controls; it does not make every destination valid or guarantee that dynamically generated content is present in the response. Store the source page and extraction policy so that missing or changed links can be diagnosed later.
Best Value
Frequently Asked Questions
Does an href always contain a complete URL?
No. It may be relative to the document, a fragment, or a value using a non-HTTP scheme. Resolve relative values against the effective base URL when needed.
Can Beautiful Soup find links created by JavaScript?
Only if those links are present in the HTML you give it. Static parsing does not itself run page JavaScript.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Should I remove URL fragments?
Only if they are irrelevant to your use case. They can identify sections within a document, so retaining them separately is often safer than discarding them.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

