Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a page whose links are present in its original HTML, fetch the document, parse every <a href>, resolve relative addresses with urljoin(), collect mailto: targets, and scan visible text for email-shaped strings. If JavaScript inserts content after load, a normal HTTP request will miss it; use a rendered browser DOM or a rendering-capable service instead. No extractor can guarantee literally every address: obfuscated text, images, inaccessible pages and non-anchor controls require separate handling.

What “all links and email addresses” should mean

Define the output before writing code. A link may be an absolute URL (https://example.com/a), a root-relative path (/about), a page-relative path (team), a fragment (#contact) or a non-navigation scheme such as javascript: or tel:. You can preserve the raw href, convert navigable values to absolute URLs, or do both.

Email addresses occur in mailto: links and ordinary text. A mailto: target can contain a query string such as ?subject=Hello; the address is the part before that query. Text may also contain false positives, so a regular expression is a practical filter, not proof that an address is valid or deliverable.

Google’s crawlable-link guidance says it can generally follow an <a> element with an href, including relative paths, but not controls that rely only on script events: Google Search Central link guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the Python dependencies

The example uses Python’s standard-library urllib.request for retrieval and Beautiful Soup for parsing. Python documents the request API at docs.python.org/3.14/library/urllib.request.html. Install Beautiful Soup with:

python -m pip install beautifulsoup4

Beautiful Soup’s built-in html.parser requires no extra parser package. Its documentation describes lxml as fast but dependent on an external C library, and html5lib as lenient and browser-like but slower. Invalid HTML can produce different trees with different parsers, so choose deliberately: Beautiful Soup 4.14.3 documentation.

Complete extractor for static HTML

Save this as extract_contacts.py, replace page_url, and run it with Python 3. The script keeps raw values, creates absolute URLs, removes fragments from the deduplicated navigation list, and gathers both mail links and visible-text matches.

from urllib.request import Request, urlopen
from urllib.parse import urljoin, urldefrag
from bs4 import BeautifulSoup
import re

page_url = "https://example.com/contact"

request = Request(
    page_url,
    headers={"User-Agent": "Mozilla/5.0 (compatible; LinkEmailExtractor/1.0)"},
)
with urlopen(request, timeout=30) as response:
    raw_html = response.read()
    # Prefer the server-provided charset; fall back only when it is absent.
    charset = response.headers.get_content_charset() or "utf-8"
    html = raw_html.decode(charset, errors="replace")

soup = BeautifulSoup(html, "html.parser")

raw_links = []
absolute_links = []
for anchor in soup.find_all("a", href=True):
    raw_href = anchor["href"].strip()
    raw_links.append(raw_href)
    absolute, _fragment = urldefrag(urljoin(page_url, raw_href))
    if absolute.startswith(("http://", "https://")):
        absolute_links.append(absolute)

# Keep first-seen order while removing duplicates.
absolute_links = list(dict.fromkeys(absolute_links))

emails = set()
for anchor in soup.find_all("a", href=True):
    href = anchor["href"].strip()
    if href.lower().startswith("mailto:"):
        address = href[len("mailto:"):].split("?", 1)[0].strip()
        if address:
            emails.add(address)

email_pattern = re.compile(
    r"[A-Za-z0-9.!#$%&'*+/=?^_`{|}~-]+@"
    r"[A-Za-z0-9-]+(?:.[A-Za-z0-9-]+)+"
)
visible_text = soup.get_text(" ", strip=True)
emails.update(email_pattern.findall(visible_text))

print("Absolute links:")
for link in absolute_links:
    print(link)
print("nEmails:")
for email in sorted(emails, key=str.casefold):
    print(email)

The response is decoded using its declared charset when available. Assuming UTF-8 for every site can corrupt non-ASCII text and cause missed matches. errors="replace" keeps the extraction running while making damaged characters visible rather than silently dropping the whole page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Raw versus normalized links

raw_links preserves exactly what authors placed in href. absolute_links applies urljoin() against the page URL, removes fragments, limits output to HTTP(S), and deduplicates while retaining first-seen order. Do not remove query strings or trailing slashes unless your application has an explicit canonicalization policy: those can identify different resources.

Case and email normalization

The script deduplicates URLs byte-for-byte after fragment removal. If your policy treats host names as case-insensitive or considers a trailing slash equivalent, normalize in a separate step and document it. For email addresses, preserve the captured spelling unless you have a reason to lowercase it; mailbox local-part handling can be more nuanced than a simple string rule.

Handling JavaScript-rendered pages

A direct fetch sees the server response, not necessarily the DOM a browser builds later. Single-page applications may add navigation, contact details or menu items only after JavaScript executes. Run a browser (for example, an automation runtime) and extract from the final DOM after the target selector appears, or use a hosted service that supports prerendering. Microlink documents link and email extraction, absolute and deduplicated URLs, and optional browser prerendering at microlink.io/use-cases/scraping/links-and-emails. Its behavior is vendor-reported, so verify current terms and limits for your workload.

In a browser workflow, wait for a meaningful condition rather than an arbitrary short delay: a selector containing the contact list, a completed network-idle state, or an application-specific “loaded” marker. Then read document.querySelectorAll('a[href]') and the rendered text. A browser still cannot reveal an address that is hidden in an image, encoded in a canvas, obfuscated as “name [at] domain [dot] com,” blocked by authentication, or never made available to the session.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the right extraction approach

Approach Best fit Trade-offs
HTTP fetch plus Beautiful Soup One or many pages whose links and text are in returned HTML Small, scriptable and inexpensive; misses client-rendered content and depends on HTTP access. Parser choice affects malformed markup.
Browser rendering or a rendering-capable service Pages that assemble links or emails with JavaScript Sees post-render DOM, but adds browser/runtime or service complexity and may require waiting, authentication and resource controls.

Decide using five questions: Is the content static? Do you need absolute and deduplicated URLs? Should prose emails count as well as mailto:? How tolerant must parsing be of malformed HTML? Is a local process appropriate for the volume and access requirements?

Or skip the browser setup

ScreenshotNeo can render a page before capture, which is useful when you need to inspect what a browser displays before running extraction. It accepts cookie and consent banners, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each cleanup step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

ScreenshotNeo returns PNG, JPEG, WebP or PDF from one GET request. The API itself is a screenshot endpoint, so you would still parse HTML or rendered DOM for links and emails; use its rendering and page-information capabilities when a plain HTTP fetch is insufficient. Every plan includes its features: 1,000 shots per month free with no card, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free.

See the ScreenshotNeo documentation for request options and authentication. A direct cURL capture looks like this:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Sign up for 1,000 free screenshots a month with no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

403, 401 or other access errors

The site may require authentication, reject your user agent, enforce rate limits or block your network. Check the response status and headers, use credentials only when authorized, slow requests, and respect the site’s terms and robots guidance. Do not assume a public URL grants permission to collect or reuse its contents.

Timeouts and incomplete output

Increase the HTTP timeout only after checking connectivity. A slow server, redirect loop or resource that never finishes can still fail. For rendered pages, wait for the actual content selector and set a maximum overall runtime. Record the final URL after redirects so relative links resolve against the document you actually received.

No emails found

Check both mailto: anchors and visible text. Inspect the raw HTML for encoded entities, then inspect the rendered DOM if JavaScript inserts the address. Images, obfuscation and protected content require site-specific handling; the regular expression cannot recover information that is not text.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Unexpected or duplicate URLs

Log raw href values alongside normalized output. Fragments, tracking parameters, alternate schemes and trailing slashes may be intentional. Define canonicalization rules before deduplicating, and never discard query parameters merely because they look unfamiliar.

Different results with different parsers

Malformed markup is repaired differently by html.parser, lxml and html5lib. Pin the parser in deployment, test representative pages, and choose the parser whose speed, dependency footprint and browser-like error recovery match your needs.

Performance, reliability and responsible use

For a single page, parsing is usually cheap compared with downloading or rendering it. Reuse an HTTP session for multiple requests, stream or cap unusually large responses, set timeouts, retry only transient failures with backoff, and cache pages when freshness permits. Keep extraction deterministic by recording the fetch time, final URL, status, parser and normalization policy.

For larger crawls, limit concurrency per host, honor access controls, avoid collecting more personal data than necessary, and protect stored addresses. The available documentation does not establish jurisdiction-specific rules for harvesting or marketing; consult applicable law, the site’s terms and your organization’s privacy guidance before bulk collection or outreach.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What this method cannot promise

  • It cannot find links that are not represented in accessible HTML or a rendered DOM.
  • It cannot reliably interpret every JavaScript event, image-embedded URL or canvas drawing.
  • It cannot prove that an email exists, accepts mail or belongs to the person named on a page.
  • It cannot bypass authentication, bot challenges or technical controls you are not authorized to pass.

Frequently Asked Questions

Should I include fragment-only links such as #pricing?

Keep them in the raw output if page navigation matters. Remove fragments only when producing deduplicated document URLs, because fragments identify a location within the same resource rather than a separate fetch target.

Why does urljoin() sometimes produce an unexpected domain?

It follows URL resolution rules: an href beginning with // inherits the current scheme, while an absolute href replaces the page’s host. Log the raw value and the page URL whenever a result is surprising.

Can a regex validate an email address completely?

No. It is a useful discovery filter. Syntax, DNS, mailbox existence and delivery policy require separate validation, and aggressive patterns can reject valid international or unusual addresses.

The Bottom Line

Use a fetch-and-parse script for links and emails already in static HTML; switch to a rendered DOM when JavaScript creates the content. Keep raw and normalized outputs separate, treat email matches as candidates, and make your scope and permission rules explicit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.