You can find email addresses exposed in a website’s public mailto: links, but public visibility is not permission to collect or use them for any purpose. For a limited, authorized task—such as checking contact links on a site you manage—the safest starting point is to inspect the specific page, collect only what you need, and record where and when you found it. Before automating anything, check the site’s terms, its robots.txt, the relevant privacy rules, and whether the site objects to scraping.
What scraping an email address means
A scraper is a program that fetches web content and extracts data that matches a pattern or location. Email addresses may appear as ordinary text, in page source, or in a link such as mailto:[email protected]. RFC 6068 warns that “’mailto’ URIs on public Web pages expose mail addresses for harvesting.” It also notes that addresses may be exposed in URI fields beyond the visible “To” field.
That is a description of a technical exposure, not permission to harvest the address. A public page can be intended for a person to contact a business while not inviting bulk collection, resale, profiling, or unsolicited marketing. The purpose and later use matter.
Check permission and scope before collecting
Start with a narrow, legitimate purpose
Write down why you need the address and what you will do with it. A one-time audit of contact links on a website you administer is different from building a prospect list from unrelated sites. If you cannot explain why each address is necessary, reduce the collection or do not proceed.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Read the site’s rules, but do not treat them as a permit
Check the site’s terms and its robots.txt file before automated access. RFC 9309 describes the Robots Exclusion Protocol as crawler guidance and states: “These rules are not a form of access authorization.” A permissive robots.txt therefore does not override terms, privacy obligations, or an explicit objection. Conversely, a restriction is a reason to stop or seek permission, not a challenge to work around.
Do not bypass a login, CAPTCHA, rate limit, or other access control to collect addresses. CNIL’s guidance on legitimate-interest analysis says that, in its context, a website’s explicit opposition to scraping through technical measures such as robots.txt or a CAPTCHA may mean collection does not meet people’s reasonable expectations. That is CNIL guidance, not a universal rule for every jurisdiction or purpose.
Decide which law applies
There is no single worldwide legal answer. Your location, the site and people involved, your purpose, and the way you process the data can all matter. The points below address EU and US sources; they are not a substitute for advice about a particular project.
Rank #2
- EU-facing processing: The European Commission identifies an email address as an example of personal data. The European Data Protection Board (EDPB) says the GDPR applies to scraping when it involves processing personal data, including collection and retrieval. Public availability alone does not remove those considerations. A lawful-basis conclusion depends on details such as the purpose, controller, people affected, and processing; do not assume that a publicly displayed address makes your project lawful.
- United States: The FTC’s CAN-SPAM guide identifies harvesting email addresses and dictionary attacks as aggravated conduct that may lead to criminal penalties, and notes that violations can carry civil penalties. This is not a claim that merely viewing or collecting every publicly displayed address invariably violates CAN-SPAM. The conduct and applicable law matter.
A limited method for checking mailto links on a page you control
The following workflow is deliberately limited to one page you own or are authorized to inspect. It reads a saved HTML file and lists the destinations of links whose href begins with mailto:. It does not crawl a site, search arbitrary text for email-like strings, evade access controls, or establish that collecting or contacting anyone is permitted.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute- Save the authorized page as HTML. Use your browser’s “Save page” option on the page you administer or have permission to inspect. Keep the saved file in a project folder and note the page URL and collection date separately.
- Save this script as
list_mailto.py. It uses only Python’s standard library and reads the local file rather than requesting pages over the network. - Run it with the file path. For example:
python3 list_mailto.py ./contact-page.html. On Windows, usepy list_mailto.py .contact-page.html. - Review results manually. Confirm that each link belongs in your audit and keep only the information your stated purpose requires.
from html.parser import HTMLParser
from pathlib import Path
from urllib.parse import urlsplit
import sys
class MailtoParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag.lower() != "a":
return
href = dict(attrs).get("href", "")
if href.lower().startswith("mailto:"):
self.links.append(href)
if len(sys.argv) != 2:
raise SystemExit("Usage: python3 list_mailto.py path/to/saved-page.html")
html = Path(sys.argv[1]).read_text(encoding="utf-8", errors="replace")
parser = MailtoParser()
parser.feed(html)
for href in parser.links:
# Display the mailto target before any query fields, such as subject or cc.
target = urlsplit(href).path
print(target)
if not parser.links:
print("No mailto links found in this saved HTML file.")
This is a page-audit aid, not a general-purpose email harvester. It will not find addresses rendered only after scripts run, embedded in images, or displayed as plain text; it can also encounter malformed or duplicate links. Those limitations are reasons to verify results against the page, not to expand the script into an indiscriminate crawler. The underlying standards establish that public mailto: links can expose addresses; they do not establish a universally reliable extraction recipe for every site.
Handle collected data carefully
Minimise and document
Keep only addresses necessary for the defined task. Maintain a record of the source page and collection time where that is appropriate to your project. The EDPB highlights purpose limitation and transparency in its scraping guidance announcement and, in the context it discusses, recommends attention to reliable sources, timestamps, validation, and data minimisation. The precise measures depend on the case.
Rank #3
Validate without turning collection into outreach
A syntactically plausible address may be outdated, mistyped, or no longer controlled by the intended person. Check it against the original authorized source and retain the source context needed for your task. Do not infer consent to contact from the presence of a link, and do not send test or marketing messages merely to see whether an address responds.
Secure, retain, and delete deliberately
Limit access to the output file, avoid placing it in public repositories or shared logs, and decide in advance when it will be deleted. If the purpose ends, remove data that is no longer needed. If a person or site objects, pause and assess the request against your purpose, legal obligations, and the applicable rules.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server—not an email extraction tool. Its one-call request returns an image or PDF of a page; it does not return email addresses or replace the limited HTML audit above. If a visual record of an authorized page is useful, the API can capture it:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Other supported client examples are:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each of those steps can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; responses report the page verdict and billing status in headers.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdffor AI agents and MCP clients. - The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Features are available on every plan.
For visual capture rather than email extraction, visit ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting the limited page audit
The script reports no links
Check that the input is the saved HTML file rather than a PDF or a screenshot, and that the page uses actual mailto: links. An address displayed as plain text will not match this script. If the page is generated dynamically, the saved source may not include the rendered link; for a site you control, inspect the rendered page and use an authorized export or audit process rather than broadening collection across other sites.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
The output contains odd text or duplicates
Review the original link in the saved HTML. A page can repeat the same contact link or encode it in an unexpected way. The script reports link targets as written before query fields; it does not normalize, verify, or deduplicate them. Do not treat a formatting cleanup as evidence that an address is valid.
You encounter a CAPTCHA, block, or explicit objection
Stop the automated access and seek permission or use an approved channel. Do not rotate identities, proxies, or accounts to continue. A crawler rule is not authorization, and a technical objection should not be treated as an invitation to bypass it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




