You can build a useful email database from public web pages, but the defensible result is a small, purpose-limited directory with evidence—not every address a crawler can discover. Define the audience and intended use first, collect only relevant contact details, record exactly where and when each address appeared, and assess the legal basis for storage and for any later message separately. A public address is not automatically permission to send marketing.
Start with a purpose, audience and jurisdiction
Write a short collection specification before opening a browser or running a crawler. It should name:
- the organizations and professional roles that qualify;
- the communication you may send (for example, a service-related introduction rather than a general newsletter);
- the countries in scope;
- the source types you will accept; and
- the conditions that disqualify an address, such as a nearby “no marketing” notice.
This is both a data-quality control and a privacy safeguard. The European Commission’s GDPR principles require specified purposes, data minimisation, accuracy and lawful, transparent processing. The UK Information Commissioner’s Office (ICO) likewise says that a public posting does not make a later direct-marketing use automatically fair or expected.
Decide whether a record identifies a natural person or only a legal entity. An address such as [email protected] identifies an employee and is personal data under EU and UK guidance even when it is displayed on a company website. A generic address such as [email protected] may still be relevant to a business purpose, but it needs a documented reason for inclusion and the same suppression controls.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Choose sources that provide professional context
Prefer pages where the organization deliberately publishes contact details for a work function: an official contact page, staff directory, investor-relations page, procurement portal or event speaker page. The ICO lists company websites, Companies House, social media and press articles as examples of public sources, while warning that personal-data obligations still apply.
A professional-network profile can show a person’s role, but using a profile for outreach may not qualify as ordinary business-to-business marketing. The ICO warns that such contact can remain subject to UK GDPR and the Privacy and Electronic Communications Regulations (PECR).
Do not treat search-engine snippets, data-broker exports or an address guessed from a naming pattern as equivalent to a published address. Canadian privacy guidance specifically notes that generating likely addresses does not create consent and warns about harvested lists.
Design an auditable record before collecting
Use a database or spreadsheet with one row per address and a stable record ID. Keep only fields tied to the stated purpose. The following schema is a practical operating standard; not every field is mandated in every country.
| Field | What to store | Why it matters |
|---|---|---|
| Organization | Legal or trading name as displayed | Separates similarly named entities and supports relevance checks. |
| Displayed name and role | Exact text on the page | Shows whether the person is acting in a professional capacity. |
| Email address | Original value, plus a normalized comparison value | Preserves evidence while allowing duplicate detection. |
| Source URL | Full page address | Lets an auditor revisit the publication context. |
| Capture date and last-checked date | UTC timestamps | Shows when the evidence was current. |
| Publication context | Nearby heading, role description and any restriction | Supports a relevance and fairness assessment. |
| Jurisdiction | Country of the person or organization, with uncertainty noted | Determines which marketing rules need review. |
| Collection assessment | Why storage appears proportionate and lawful | Keeps collection and later sending decisions distinct. |
| Campaign relevance | Why the planned message relates to the role | Important where a narrow business-role exception is considered. |
| Notice and suppression status | Objection, unsubscribe, do-not-contact flag and date | Prevents a later campaign from reviving a rejected address. |
Retain a short text excerpt or a screenshot of the address and its surrounding context when justified by your retention policy. Do not store an entire page merely because it is available.
DIY collection workflow
1. Set crawling limits
Use a slow, identifiable client, follow the site’s terms and access controls, honor applicable robots instructions, and avoid login-only areas. Start with a small list of known domains instead of searching the whole web. A public page can still contain personal data, and technical access does not establish marketing permission.
Rank #2
2. Extract only visibly published addresses
The following Python example reads a list of pages, extracts visible email links and ordinary email text, and writes provenance fields. It is intentionally conservative: it does not guess addresses, bypass access controls or send messages.
import csv
import re
import time
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
EMAIL_RE = re.compile(r"[A-Z0-9._%+-]+@[A-Z0-9.-]+.[A-Z]{2,}", re.I)
HEADERS = {"User-Agent": "PublicContactResearch/1.0 (contact: [email protected])"}
def collect_page(url):
response = requests.get(url, headers=HEADERS, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
text = soup.get_text(" ", strip=True)
found = set(EMAIL_RE.findall(text))
for link in soup.select('a[href^="mailto:"]'):
address = link.get("href", "")[7:].split("?", 1)[0]
if EMAIL_RE.fullmatch(address):
found.add(address)
title = soup.title.get_text(" ", strip=True) if soup.title else ""
captured = datetime.now(timezone.utc).isoformat()
return [{
"email_original": address,
"email_normalized": address.casefold(),
"source_url": url,
"page_title": title,
"captured_at": captured,
"publication_context": "Visible page text or mailto link; review manually",
} for address in sorted(found, key=str.casefold)]
with open("source_urls.txt", encoding="utf-8") as source_file,
open("contacts_raw.csv", "w", newline="", encoding="utf-8") as output_file:
writer = csv.DictWriter(output_file, fieldnames=[
"email_original", "email_normalized", "source_url", "page_title",
"captured_at", "publication_context"
])
writer.writeheader()
for line in source_file:
url = line.strip()
if not url or url.startswith("#"):
continue
try:
for row in collect_page(url):
writer.writerow(row)
except requests.RequestException as error:
print(f"Skipped {url}: {error}")
time.sleep(2) # keep request volume modest
Install the two dependencies with python -m pip install requests beautifulsoup4. Put one authorized, public URL per line in source_urls.txt. The output is a review queue, not a send-ready list.
3. Review every match
Automated extraction can capture an example address in a privacy policy, an image alt text, a code sample or a third party’s quotation. Open the source page and confirm that the address is actually published for the professional context you recorded. Note any “do not contact,” “no unsolicited email” or similar language. Remove addresses that fail your audience definition.
4. Normalize and deduplicate without destroying evidence
Use a case-folded comparison value for ordinary duplicate detection, but retain the original spelling. Do not apply provider-specific transformations such as removing dots or plus-tags unless you have a documented, provider-specific reason; those changes can merge different mailboxes. If the same address appears on several pages, keep each source as a provenance entry or maintain a linked source table.
5. Separate collection permission from sending permission
For each surviving record, write two decisions: why you may retain the data, and why the proposed message may be sent. A lawful basis for maintaining a directory does not automatically authorize a commercial email campaign.
- United Kingdom: The ICO says public availability is not assumed consent. UK GDPR duties apply to publicly available personal data, and PECR adds rules for electronic marketing.
- European Union: GDPR covers identifiable people acting professionally. Data about a company as a legal entity alone is outside GDPR, but an employee’s named address is not. Direct-marketing email also engages applicable ePrivacy rules.
- Canada: CASL generally requires express or qualifying implied consent. CRTC guidance allows a narrow conspicuous-publication route only when no contrary statement appears and the message relates to the recipient’s business role, functions or duties. The sender must prove those conditions.
- United States: The FTC says CAN-SPAM covers commercial email, including business-to-business messages. Headers and subject lines must be accurate, the message needs a physical postal address and an opt-out method, and opt-outs must be honored within 10 business days.
Compare source and contact types before inclusion
| Source or contact | Useful evidence | Questions to resolve |
|---|---|---|
| Official company contact page | Organization, function and publication context | Is the address generic or tied to an identifiable employee? Is there a no-contact statement? |
| Staff or leadership page | Name, role and business purpose | Does the planned message genuinely relate to that role? |
| Government or corporate registry | Entity identity and official filings | Does the page publish an email for contact, and which jurisdiction governs use? |
| Professional-network profile | Current role and professional capacity | Would outreach be treated as personal-data direct marketing rather than ordinary B2B contact? |
| Press article or event page | Time-stamped context | Is the address current, and was it posted for media or event logistics rather than solicitation? |
| Guessed or generated address | None proving publication or consent | Exclude it; pattern matching does not supply permission. |
Assess every source against seven axes: person versus legal entity, publication context, any explicit restriction, contemporaneous proof, role relevance, jurisdiction and channel, and whether objections can be propagated to every copy.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
Maintain objections, accuracy and vendor controls
Attach evidence to the record instead of accepting a supplier’s statement that a list is compliant. For a Canadian conspicuous-publication assessment, CRTC guidance recommends contemporaneous proof of the address, date, URL, absence of a no-message instruction and business relevance. The FTC requires prompt opt-out handling and restricts later transfer of opted-out addresses except to a compliance service provider.
Keep a suppression table separate from your active prospects. When someone objects, record the address, date, channel and source of the objection, then block it in every campaign system and backup export. Do not delete the only suppression record: deletion can cause the address to be re-imported later.
A contractor does not remove your responsibility. Canadian Office of the Privacy Commissioner guidance says an organization remains accountable when a supplier provides a list or runs a campaign. Require suppliers to explain collection sources, update procedures, suppression propagation and how they handle withdrawn consent.
Before each campaign, verify that the address still works, the person’s role and relevance are current, no objection is recorded, and the message remains within the documented purpose and jurisdictional rule. Set a risk-based review cadence for dormant records and record the decision.
Recommended Free Tools
Performance, reliability and cost controls
- Cache page results and avoid recrawling unchanged domains; caching reduces load and makes an audit trail easier to reproduce.
- Use bounded timeouts, retries with backoff and a failed-URL queue. A timeout is a collection failure, not evidence that an address is absent.
- Log HTTP status, redirect destination and capture time. Treat bot checks, blank responses and consent overlays as review states rather than successful extraction.
- Keep extraction and outreach as separate jobs. A parser should write records; a human or controlled compliance step should authorize any campaign import.
- Budget for manual review. The expensive part of a defensible database is usually verifying context, jurisdiction and objections, not finding a string that resembles an email address.
Troubleshooting common failures
The script finds no addresses
The page may render email content with JavaScript, use an image, require a consent interaction or return a bot-check page. Open the page manually, record the failure state, and use an approved browser workflow only when the site permits it. Do not defeat a CAPTCHA or access control.
Every page returns 403 or 429
Slow the request rate, identify the client, stop when the site signals blocking and ask the site owner for an approved export or access method. Repeated retries can worsen the block and do not create a right to continue.
One address appears under several people
It may be a role inbox or a forwarding address. Preserve each publication context, classify the inbox correctly and avoid attaching a named person unless the page does so.
A contact asks where you obtained the address
Provide the recorded source page and capture date, explain the intended purpose in plain language, and offer an immediate removal route. Record the request in the suppression table.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
A vendor’s list has no source URLs
Do not import it into an active campaign. Request row-level provenance and suppression history; without evidence, you cannot test relevance, publication conditions or objections.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo can capture an approved public page for your evidence record through one request. It is a screenshot API and MCP server, not a permission system: a screenshot does not make an address lawful to collect or market to. Before capture, it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
Use the API documentation at https://screenshotneo.com/docs/ for options such as full-page capture with lazy images loaded, CSS-element capture, custom JavaScript and CSS, waits, blocked requests, headers and cookies, timezone and geolocation, PDF output, caching with a chosen TTL, signed links and asynchronous webhooks. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing provides two months free. Create a free ScreenshotNeo account to capture source evidence without installing a browser.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsFAQ
Does robots.txt grant permission to send marketing?
No. Robots instructions concern automated access; they do not establish consent, a lawful basis or compliance with electronic-marketing rules.
Best Value
How should I handle a page that disappears?
Keep the original URL, capture date and retained context, mark the source as unavailable, and do not treat the old publication as proof that the address remains current. Re-verify the contact before any campaign.
Can a screenshot prove consent?
No. It can preserve what a page displayed, including a restriction or role description. Consent, fairness and sending permission require a separate documented assessment.
Frequently Asked Questions
Does robots.txt grant permission to send marketing?
No. Robots instructions concern automated access; they do not establish consent, a lawful basis or compliance with electronic-marketing rules.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →How should I handle a page that disappears?
Keep the original URL, capture date and retained context, mark the source as unavailable, and re-verify the contact before any campaign.
Can a screenshot prove consent?
No. It preserves what a page displayed, but consent, fairness and sending permission require a separate documented assessment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




