Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The reliable way to extract links is to fetch a page, select its <a> elements, read each href, and resolve relative URLs against the page URL (or the document’s <base> element). For a whole site, add a queue, domain and URL rules, deduplication, crawl limits, and robots.txt handling. If links are added by JavaScript, inspect the browser’s network requests or render the page with a headless browser instead of assuming the visible browser page matches the original HTML.

Choose the extraction method first

Your method depends on four questions:

  • Are you extracting one page or recursively crawling a site?
  • Are the links present in the initial HTML response, or loaded later by JavaScript?
  • Do you need filtering by domain, path, CSS/XPath region, tag, or URL pattern?
  • What page, depth, request-rate, duplicate, and robots.txt limits apply?

For a one-off job, an HTTP client and HTML parser are usually simplest. For a controlled crawl, Scrapy supplies scheduling, duplicate filtering, and link-extraction rules. For client-rendered pages, reproduce the data request found in browser developer tools or use a headless browser when the content is available only in the rendered DOM.

Extract links from one page with Python

Install the parser and HTTP client

python -m pip install requests beautifulsoup4

Complete, runnable example

from urllib.parse import urljoin, urldefrag
import requests
from bs4 import BeautifulSoup

page_url = "https://example.com/"
response = requests.get(
    page_url,
    headers={"User-Agent": "link-extractor/1.0"},
    timeout=30,
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
base_tag = soup.find("base", href=True)
base_url = urljoin(response.url, base_tag["href"]) if base_tag else response.url

seen = set()
for anchor in soup.select("a[href]"):
    raw_href = anchor["href"].strip()
    if not raw_href:
        continue
    absolute = urljoin(base_url, raw_href)
    absolute, fragment = urldefrag(absolute)
    if absolute in seen:
        continue
    seen.add(absolute)
    text = " ".join(anchor.get_text(" ", strip=True).split())
    print({"url": absolute, "text": text, "fragment": fragment or None})

This preserves the visible link text, removes URL fragments for deduplication, and records the fragment separately. urljoin handles paths such as /docs, ../pricing, and query strings correctly. It also prevents the common mistake of treating every href as an already absolute URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle non-navigational href values

Not every href is a page to request. Skip or classify values such as mailto:, tel:, javascript:, and data URLs when your output is intended to contain web pages only:

from urllib.parse import urlparse

scheme = urlparse(absolute).scheme.lower()
if scheme not in {"http", "https"}:
    continue

Keep them in a separate output if you are auditing every hyperlink rather than building a crawl queue.

Filter to the same site

from urllib.parse import urlparse

site_host = urlparse(response.url).netloc.lower()
link_host = urlparse(absolute).netloc.lower()
if link_host != site_host:
    continue

Decide whether subdomains are in scope. Exact host matching treats www.example.com and blog.example.com as different hosts; that is safer than silently crawling every subdomain.

Extract only links in a page region

CSS selectors let you restrict extraction to navigation, an article, or a card list:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
for anchor in soup.select("main article a[href]"):
    print(urljoin(base_url, anchor["href"]))

Use an XPath-capable parser when the page structure is easier to express that way. A region selector is important when sidebars, footers, or recommendation widgets would otherwise pollute your result.

Crawl a website with Scrapy

Extraction and following are separate decisions: first identify links worth collecting; then decide which of those links the crawler may request. Scrapy’s LxmlLinkExtractor is designed for the second task. Its defaults examine a and area tags and their href attribute. It can filter by allowed or denied URL patterns, domains, CSS or XPath regions, tags, attributes, and duplicate handling. Returned links can include the URL, text, fragment, and a nofollow indicator.

Minimal spider

import scrapy
from scrapy.linkextractors import LinkExtractor

class SiteSpider(scrapy.Spider):
    name = "site_links"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    def parse(self, response):
        extractor = LinkExtractor(
            allow=(r"/docs/",),
            deny=(r"/login", r"/logout"),
            unique=True,
        )
        for link in extractor.extract_links(response):
            yield {
                "url": link.url,
                "text": link.text,
                "fragment": link.fragment,
                "nofollow": link.nofollow,
            }
            yield response.follow(link.url, callback=self.parse)

Set an explicit start URL, allowed domains, URL patterns, and crawl limits in your project settings. Restrict depth and page count for exploratory jobs; otherwise a calendar, search, or faceted-navigation URL can create a very large crawl. Store a canonical form of each URL and keep request concurrency and delays appropriate for the site.

Resolve relative URLs correctly

Scrapy resolves relative links using an HTML <base> element when one is present and otherwise the response URL. Do not concatenate strings such as response.url + href; that breaks parent paths, query strings, and fragments.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Respect robots.txt before crawling

Check the site’s top-level /robots.txt before sending a crawl. RFC 9309, the Robots Exclusion Protocol published in September 2022, defines robots.txt as crawler guidance. It says: “These rules are not a form of access authorization.” A successfully retrieved file still contains parseable rules your crawler should follow; it does not grant permission to access restricted resources, bypass authentication, or ignore other legal and contractual limits.

  • Fetch the file from the service’s top-level path, such as https://example.com/robots.txt.
  • Apply the rules to the user agent your crawler identifies.
  • Do not treat a missing or unreachable file as permission to overload the site.
  • Keep request rates, depth, and page limits conservative and identify your crawler honestly.

When browser-visible links are missing

A normal HTTP response may contain only an application shell. The browser can then request JSON, GraphQL, or HTML fragments and insert links into the DOM. Compare the downloaded response with what you see in the browser:

  1. Open developer tools and select the Network panel.
  2. Reload the page and filter requests by Fetch/XHR (and, when relevant, document or GraphQL).
  3. Inspect responses for the URL data or HTML containing the missing anchors.
  4. Reproduce the request with the required method, query parameters, headers, cookies, or authorization, subject to the site’s rules.
  5. If reproducing requests is impractical but the content is accessible in the browser DOM, use a headless browser and extract anchors after the relevant page state appears.

Do not rely on a fixed sleep alone when possible. Wait for a selector or a network-idle condition, and record the exact state that produced the links. Some links appear only after scrolling, clicking a tab, accepting consent, or signing in; those are distinct workflows with their own permissions and failure modes.

Normalize, deduplicate, and classify results

  • Fragments: remove #section before crawl deduplication, but retain it if the destination within a document matters.
  • Trailing slashes and case: do not rewrite them blindly; servers can treat distinct paths differently.
  • Query parameters: decide which tracking or filter parameters are safe to drop. Removing a parameter can merge genuinely different pages.
  • Redirects: keep both the discovered URL and the final response URL when auditing redirects.
  • Duplicates: use a set or Scrapy’s uniqueness controls before scheduling requests.
  • Link type: classify HTTP links separately from email, telephone, JavaScript, and data links.
  • Context: retain anchor text, source page, region, fragment, and nofollow state when the output will be reviewed by people.

cURL, Python, and Node.js request patterns

For a simple HTML download, cURL is useful for checking exactly what the server returns:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -L --max-time 30 -A "link-extractor/1.0" https://example.com/ -o page.html

Python’s requests example above parses the response directly. In Node.js, fetch the HTML and pass it to an HTML parser such as Cheerio:

import * as cheerio from "cheerio";

const pageUrl = "https://example.com/";
const res = await fetch(pageUrl, {
  headers: { "user-agent": "link-extractor/1.0" }
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const html = await res.text();
const $ = cheerio.load(html);
const baseHref = $("base[href]").first().attr("href");
const baseUrl = new URL(baseHref || pageUrl, pageUrl);

const links = [];
$("a[href]").each((_, el) => {
  const raw = $(el).attr("href").trim();
  const target = new URL(raw, baseUrl);
  if (!["http:", "https:"].includes(target.protocol)) return;
  target.hash = "";
  links.push({ url: target.href, text: $(el).text().trim() });
});
console.log(links);

Or skip the browser setup

When your goal is a clean visual record of pages rather than parsing their anchors, ScreenshotNeo provides a one-request website screenshot API. It accepts a URL and returns PNG, JPEG, WebP, or PDF; its cleanup steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

Use the documented options at ScreenshotNeo’s API documentation for full-page captures, lazy-image loading, CSS-selector element capture, dark mode, device and viewport settings, retina scale, PDF paper and margin controls, custom CSS or JavaScript, click and wait conditions, blocked ads or resources, headers, cookies, user agents, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and the OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

The output is empty

Check the response status, content type, and saved HTML. You may have received a login page, consent interstitial, bot challenge, or JavaScript shell rather than the intended document. Reproduce the data request or render the page after the required state is reached.

URLs point to the wrong directory

Use urljoin or Scrapy’s resolver and account for <base href>. Never prepend the host with string operations.

The crawler never stops

Add allowed domains, allow/deny patterns, depth and page limits, and deduplication. Faceted navigation and calendar parameters often generate effectively unbounded URL spaces.

Requests are blocked or receive 403

Verify that your request is permitted, identify your client honestly, honor robots.txt, slow the crawl, and use the site’s documented API when available. Do not attempt to bypass authentication or access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links are duplicated

Normalize fragments for crawl identity, preserve query parameters unless you have a documented canonicalization rule, and deduplicate before scheduling requests.

FAQ

Can I extract links without downloading the whole site?

Yes. Fetch a single page, parse its anchors, and stop there; recursive following is an additional operation, not a requirement.

Should nofollow links be excluded?

Not automatically. Keep their nofollow state as metadata, then decide whether your particular audit or crawl should follow them.

Is robots.txt permission to scrape?

No. It provides crawler guidance, not authorization, as RFC 9309 states.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What is the fastest way to get every URL from one HTML page?

Fetch the page, select a[href], resolve each value with the response URL or <base>, and deduplicate according to your fragment and query-string policy.

Why do links visible in Chrome not appear in my Python response?

They may be injected after JavaScript runs. Inspect Fetch/XHR requests, reproduce the underlying response, or render the page and extract the post-load DOM.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.