Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The reliable way to extract links is to fetch a page, select its <a> elements, read each href, and resolve relative URLs against the page URL (or the document’s <base> element). For a whole site, add a queue, domain and URL rules, deduplication, crawl limits, and robots.txt handling. If links are added by JavaScript, inspect the browser’s network requests or render the page with a headless browser instead of assuming the visible browser page matches the original HTML.
Choose the extraction method first
Your method depends on four questions:
- Are you extracting one page or recursively crawling a site?
- Are the links present in the initial HTML response, or loaded later by JavaScript?
- Do you need filtering by domain, path, CSS/XPath region, tag, or URL pattern?
- What page, depth, request-rate, duplicate, and robots.txt limits apply?
For a one-off job, an HTTP client and HTML parser are usually simplest. For a controlled crawl, Scrapy supplies scheduling, duplicate filtering, and link-extraction rules. For client-rendered pages, reproduce the data request found in browser developer tools or use a headless browser when the content is available only in the rendered DOM.
Extract links from one page with Python
Install the parser and HTTP client
python -m pip install requests beautifulsoup4
Complete, runnable example
from urllib.parse import urljoin, urldefrag
import requests
from bs4 import BeautifulSoup
page_url = "https://example.com/"
response = requests.get(
page_url,
headers={"User-Agent": "link-extractor/1.0"},
timeout=30,
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
base_tag = soup.find("base", href=True)
base_url = urljoin(response.url, base_tag["href"]) if base_tag else response.url
seen = set()
for anchor in soup.select("a[href]"):
raw_href = anchor["href"].strip()
if not raw_href:
continue
absolute = urljoin(base_url, raw_href)
absolute, fragment = urldefrag(absolute)
if absolute in seen:
continue
seen.add(absolute)
text = " ".join(anchor.get_text(" ", strip=True).split())
print({"url": absolute, "text": text, "fragment": fragment or None})
This preserves the visible link text, removes URL fragments for deduplication, and records the fragment separately. urljoin handles paths such as /docs, ../pricing, and query strings correctly. It also prevents the common mistake of treating every href as an already absolute URL.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Handle non-navigational href values
Not every href is a page to request. Skip or classify values such as mailto:, tel:, javascript:, and data URLs when your output is intended to contain web pages only:
#1 Best Overall
from urllib.parse import urlparse
scheme = urlparse(absolute).scheme.lower()
if scheme not in {"http", "https"}:
continue
Keep them in a separate output if you are auditing every hyperlink rather than building a crawl queue.
Filter to the same site
from urllib.parse import urlparse
site_host = urlparse(response.url).netloc.lower()
link_host = urlparse(absolute).netloc.lower()
if link_host != site_host:
continue
Decide whether subdomains are in scope. Exact host matching treats www.example.com and blog.example.com as different hosts; that is safer than silently crawling every subdomain.
Extract only links in a page region
CSS selectors let you restrict extraction to navigation, an article, or a card list:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
for anchor in soup.select("main article a[href]"):
print(urljoin(base_url, anchor["href"]))
Use an XPath-capable parser when the page structure is easier to express that way. A region selector is important when sidebars, footers, or recommendation widgets would otherwise pollute your result.
Crawl a website with Scrapy
Extraction and following are separate decisions: first identify links worth collecting; then decide which of those links the crawler may request. Scrapy’s LxmlLinkExtractor is designed for the second task. Its defaults examine a and area tags and their href attribute. It can filter by allowed or denied URL patterns, domains, CSS or XPath regions, tags, attributes, and duplicate handling. Returned links can include the URL, text, fragment, and a nofollow indicator.
Rank #2
Minimal spider
import scrapy
from scrapy.linkextractors import LinkExtractor
class SiteSpider(scrapy.Spider):
name = "site_links"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/"]
def parse(self, response):
extractor = LinkExtractor(
allow=(r"/docs/",),
deny=(r"/login", r"/logout"),
unique=True,
)
for link in extractor.extract_links(response):
yield {
"url": link.url,
"text": link.text,
"fragment": link.fragment,
"nofollow": link.nofollow,
}
yield response.follow(link.url, callback=self.parse)
Set an explicit start URL, allowed domains, URL patterns, and crawl limits in your project settings. Restrict depth and page count for exploratory jobs; otherwise a calendar, search, or faceted-navigation URL can create a very large crawl. Store a canonical form of each URL and keep request concurrency and delays appropriate for the site.
Resolve relative URLs correctly
Scrapy resolves relative links using an HTML <base> element when one is present and otherwise the response URL. Do not concatenate strings such as response.url + href; that breaks parent paths, query strings, and fragments.
Respect robots.txt before crawling
Check the site’s top-level /robots.txt before sending a crawl. RFC 9309, the Robots Exclusion Protocol published in September 2022, defines robots.txt as crawler guidance. It says: “These rules are not a form of access authorization.” A successfully retrieved file still contains parseable rules your crawler should follow; it does not grant permission to access restricted resources, bypass authentication, or ignore other legal and contractual limits.
- Fetch the file from the service’s top-level path, such as
https://example.com/robots.txt. - Apply the rules to the user agent your crawler identifies.
- Do not treat a missing or unreachable file as permission to overload the site.
- Keep request rates, depth, and page limits conservative and identify your crawler honestly.
When browser-visible links are missing
A normal HTTP response may contain only an application shell. The browser can then request JSON, GraphQL, or HTML fragments and insert links into the DOM. Compare the downloaded response with what you see in the browser:
- Open developer tools and select the Network panel.
- Reload the page and filter requests by Fetch/XHR (and, when relevant, document or GraphQL).
- Inspect responses for the URL data or HTML containing the missing anchors.
- Reproduce the request with the required method, query parameters, headers, cookies, or authorization, subject to the site’s rules.
- If reproducing requests is impractical but the content is accessible in the browser DOM, use a headless browser and extract anchors after the relevant page state appears.
Do not rely on a fixed sleep alone when possible. Wait for a selector or a network-idle condition, and record the exact state that produced the links. Some links appear only after scrolling, clicking a tab, accepting consent, or signing in; those are distinct workflows with their own permissions and failure modes.
Normalize, deduplicate, and classify results
- Fragments: remove
#sectionbefore crawl deduplication, but retain it if the destination within a document matters. - Trailing slashes and case: do not rewrite them blindly; servers can treat distinct paths differently.
- Query parameters: decide which tracking or filter parameters are safe to drop. Removing a parameter can merge genuinely different pages.
- Redirects: keep both the discovered URL and the final response URL when auditing redirects.
- Duplicates: use a set or Scrapy’s uniqueness controls before scheduling requests.
- Link type: classify HTTP links separately from email, telephone, JavaScript, and data links.
- Context: retain anchor text, source page, region, fragment, and nofollow state when the output will be reviewed by people.
cURL, Python, and Node.js request patterns
For a simple HTML download, cURL is useful for checking exactly what the server returns:
curl -L --max-time 30 -A "link-extractor/1.0" https://example.com/ -o page.html
Python’s requests example above parses the response directly. In Node.js, fetch the HTML and pass it to an HTML parser such as Cheerio:
import * as cheerio from "cheerio";
const pageUrl = "https://example.com/";
const res = await fetch(pageUrl, {
headers: { "user-agent": "link-extractor/1.0" }
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const html = await res.text();
const $ = cheerio.load(html);
const baseHref = $("base[href]").first().attr("href");
const baseUrl = new URL(baseHref || pageUrl, pageUrl);
const links = [];
$("a[href]").each((_, el) => {
const raw = $(el).attr("href").trim();
const target = new URL(raw, baseUrl);
if (!["http:", "https:"].includes(target.protocol)) return;
target.hash = "";
links.push({ url: target.href, text: $(el).text().trim() });
});
console.log(links);
Or skip the browser setup
When your goal is a clean visual record of pages rather than parsing their anchors, ScreenshotNeo provides a one-request website screenshot API. It accepts a URL and returns PNG, JPEG, WebP, or PDF; its cleanup steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
Use the documented options at ScreenshotNeo’s API documentation for full-page captures, lazy-image loading, CSS-selector element capture, dark mode, device and viewport settings, retina scale, PDF paper and margin controls, custom CSS or JavaScript, click and wait conditions, blocked ads or resources, headers, cookies, user agents, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and the OpenAPI specification.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Troubleshooting common failures
The output is empty
Check the response status, content type, and saved HTML. You may have received a login page, consent interstitial, bot challenge, or JavaScript shell rather than the intended document. Reproduce the data request or render the page after the required state is reached.
URLs point to the wrong directory
Use urljoin or Scrapy’s resolver and account for <base href>. Never prepend the host with string operations.
The crawler never stops
Add allowed domains, allow/deny patterns, depth and page limits, and deduplication. Faceted navigation and calendar parameters often generate effectively unbounded URL spaces.
Requests are blocked or receive 403
Verify that your request is permitted, identify your client honestly, honor robots.txt, slow the crawl, and use the site’s documented API when available. Do not attempt to bypass authentication or access controls.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsSome links are duplicated
Normalize fragments for crawl identity, preserve query parameters unless you have a documented canonicalization rule, and deduplicate before scheduling requests.
FAQ
Can I extract links without downloading the whole site?
Yes. Fetch a single page, parse its anchors, and stop there; recursive following is an additional operation, not a requirement.
Best Value
Should nofollow links be excluded?
Not automatically. Keep their nofollow state as metadata, then decide whether your particular audit or crawl should follow them.
Is robots.txt permission to scrape?
No. It provides crawler guidance, not authorization, as RFC 9309 states.
Recommended Free Tools
Frequently Asked Questions
What is the fastest way to get every URL from one HTML page?
Fetch the page, select a[href], resolve each value with the response URL or <base>, and deduplicate according to your fragment and query-string policy.
Why do links visible in Chrome not appear in my Python response?
They may be injected after JavaScript runs. Inspect Fetch/XHR requests, reproduce the underlying response, or render the page and extract the post-load DOM.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

