October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
data quality

Company Website Scraping and Lead Enrichment: A Practical, Compliance-Aware Workflow

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrape company websites for lead enrichment only after you define a narrow purpose, the exact fields you need, and the access rules that apply to each source. Prefer an owner-provided API or data feed. If you must collect pages directly, use a low-volume crawler that respects stated objections and technical controls, stores every value with its source and collection time, removes irrelevant data quickly, and validates records before anyone uses them for outreach. Public visibility alone does not authorize every reuse, and a technically successful crawl is not proof that a campaign is lawful.

What company website scraping and lead enrichment actually involve

Website scraping is the mechanical retrieval of pages or structured responses. Lead enrichment is the later work of turning those responses into usable business records: identifying the organization, normalizing fields, checking whether values are current, and preserving enough provenance for another person to audit the result.

A defensible record normally includes the company name, canonical website, a source URL for each important value, the collection timestamp, and a freshness or review status. Treat a scraped value as an observation that needs checking, not as ground truth. The available guidance establishes no general accuracy, conversion, or return-on-investment rate for scraped leads.

Is scraping public company information legal?

There is no universal yes-or-no answer. CNIL’s web-scraping focus sheet (5 January 2026) says scraping is not prohibited per se and requires a case-by-case assessment. The same guidance warns that other rules can apply, including contractual terms, database rights, copyright, and data-protection requirements. The intended use, the source, the fields, and the people affected all matter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Publicly visible does not mean unrestricted reuse

A page that anyone can view may still carry terms, a licensing condition, an objection to automated collection, or information published in a context that does not reasonably support your planned reuse. Separate the decision to collect a fact about a company from the decision to contact an individual. A generic company address and a named employee’s direct address raise different privacy and outreach questions.

Define purpose and minimization before the first request

Write down why you are collecting data, which companies and domains are in scope, the fields required to achieve that purpose, and the retention period. CNIL’s guidance calls for specific criteria in advance, exclusion of unnecessary categories, and prompt deletion of irrelevant information. Do not collect every visible field merely because your parser can.

Take objections and access controls seriously

Review the site’s terms and its technical signals before collection. In the context described by CNIL, robots.txt and CAPTCHA can be clear signals to exclude a site; the guidance also discusses legal objections such as terms of service. Do not bypass a login wall, CAPTCHA, rate limit, paywall, or other access control. If a source says no, use an API, request permission, or remove it from the job.

Understand what robots.txt does—and does not do

Google’s robots.txt documentation describes crawler instructions for a particular protocol, host, and port. Google Search Central also explains that robots.txt is not a security boundary, does not guarantee that a URL disappears from search results, and may be interpreted differently by different crawlers. Therefore, a disallow rule is neither a universal legal ruling nor a license to ignore the site’s stated preferences. Treat it as an important technical objection and document how your crawler handled it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the GDPR discussion in context

CNIL’s recommendations for AI-system development (accessed 29 September 2026) state that web scraping is not, in itself, prohibited under the GDPR and that a private body may rely on legitimate interest when it implements appropriate safeguards. That statement is framed around AI-system development, not a blanket approval for lead generation. For an enrichment project, assess lawful basis, transparency, objections, retention, and the direct-marketing rules that apply to the recipients’ jurisdictions. Obtain advice for high-risk or cross-border campaigns.

A bounded workflow from source selection to usable leads

  1. Write a collection specification. List the business question, target domains, required fields, exclusion rules, refresh interval, retention limit, and who may use the output.
  2. Assess each source. Read terms, inspect robots.txt, note login requirements, CAPTCHAs, rate limits, and any owner contact or API documentation. Record the date of the assessment.
  3. Try an authorized route first. Ask for an API, export, or direct feed. Eurostat’s Practical guidelines on web scraping for the HICP (November 2020) describe APIs as structured access and recommend contacting site owners. Its discussion of third-party tools is specific to statistical collection, so use it as an engineering comparison rather than a legal conclusion.
  4. Collect the minimum fields. Keep company facts separate from personal data. Exclude sensitive or irrelevant information unless it is necessary, assessed, and covered by your process.
  5. Preserve provenance. Store the exact source URL, retrieval time in UTC, parser version, response status, and a hash or snapshot reference when your retention policy permits.
  6. Validate before enrichment. Check required formats, domain ownership, duplicate organizations, stale pages, and contradictory values. Route uncertain records to a review queue instead of silently guessing.
  7. Review downstream use. A record can be technically valid and still unsuitable for outreach. Apply the rules for legal basis, notice, opt-out handling, suppression lists, retention, and the recipient’s country before sending anything.

Choose the collection route deliberately

Route Use it when Questions to document
Source-provided API The publisher exposes the fields you need in a structured interface. Is access authorized? What are field coverage, stability, update cadence, rate limits, and cost?
Owner access or data feed The data is important, recurring, or likely to break without coordination. What permission, refresh schedule, support, change notices, completeness, and fees apply?
Team-operated scraper The source permits collection and you need control over scope and processing. How will you handle source changes, request volume, validation, maintenance, and audit logs?
Third-party scraping application No suitable API exists and a managed tool fits the source and workflow. What are the charges, script controls, source restrictions, storage location, integrations, and export terms?

Eurostat notes that third-party applications can involve charges, limits on script changes, and implications for where data is stored. Check current terms and features directly; no vendor ranking or affiliate relationship is established here.

Design a lead-enrichment record that can be audited

Use a stable company key rather than the page title alone. A practical schema is:

  • Identity: legal or trading name as displayed, normalized domain, country, and an internal organization ID.
  • Business facts: description, products or services, industries, locations, and public contact channels that are necessary for the stated purpose.
  • Evidence: source URL per field, collected_at timestamp, HTTP status, parser version, and a short evidence note.
  • Quality state: unreviewed, validated, conflicting, stale, excluded, or deleted.
  • Governance: purpose, retention deadline, objection status, suppression status, and reviewer.

Normalize domains consistently, but retain the original URL. De-duplicate on more than a company name: compare canonical domains, legal names, addresses, and country. If two pages disagree, keep both observations with timestamps and send the record for review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A conservative do-it-yourself scraper in Python

The example below collects only a page title, meta description, canonical URL, and Organization JSON-LD. It checks robots.txt as an operational signal, uses a descriptive user agent, spaces requests, limits the URL list, and writes provenance. It does not defeat access controls or attempt to harvest personal email addresses.

import json
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup

USER_AGENT = "ExampleLeadResearchBot/1.0 (contact: [email protected])"
URLS = [
    "https://example.com/",
]

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})


def robots_allows(url):
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    parser = RobotFileParser(robots_url)
    try:
        parser.read()
        return parser.can_fetch(USER_AGENT, url), robots_url
    except Exception:
        # A fetch failure is a reason to pause and review, not to bypass controls.
        return False, robots_url


def extract_record(url):
    allowed, robots_url = robots_allows(url)
    if not allowed:
        return {"url": url, "status": "excluded", "reason": "robots policy or unavailable robots.txt", "robots_url": robots_url}

    response = session.get(url, timeout=30)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    title = soup.title.get_text(" ", strip=True) if soup.title else None
    description_tag = soup.find("meta", attrs={"name": "description"})
    canonical_tag = soup.find("link", rel="canonical")
    organizations = []
    for node in soup.find_all("script", attrs={"type": "application/ld+json"}):
        try:
            data = json.loads(node.string or "")
            items = data if isinstance(data, list) else [data]
            for item in items:
                if isinstance(item, dict) and item.get("@type") in ("Organization", "Corporation", "LocalBusiness"):
                    organizations.append(item)
        except (json.JSONDecodeError, TypeError):
            continue

    return {
        "url": url,
        "canonical_url": urljoin(url, canonical_tag.get("href")) if canonical_tag else None,
        "title": title,
        "description": description_tag.get("content", "").strip() if description_tag else None,
        "organization_jsonld": organizations,
        "collected_at": datetime.now(timezone.utc).isoformat(),
        "http_status": response.status_code,
        "status": "collected",
    }


records = []
for target in URLS:
    try:
        records.append(extract_record(target))
    except requests.RequestException as exc:
        records.append({"url": target, "status": "failed", "error": str(exc)})
    time.sleep(2)  # Set a rate that your source permits; do not assume this is universal.

with open("company_records.json", "w", encoding="utf-8") as output:
    json.dump(records, output, ensure_ascii=False, indent=2)

Install the two dependencies in an isolated environment with python -m pip install requests beautifulsoup4. Before production use, replace the example domain, obtain approval for the source list, add a maximum-page and maximum-byte budget, and test how your parser handles redirects, compressed responses, non-HTML content, and consent pages. A robots parser is only one implementation of a convention; it is not a legal determination and should not be treated as permission to crawl a site that otherwise objects.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When a page must be rendered to verify what a visitor sees, ScreenshotNeo is a practical rendering layer rather than a substitute for permission to collect data. It accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Use the API documentation at https://screenshotneo.com/docs/. A one-call example is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For automated workflows, its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Other useful controls include full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper size, margins, landscape mode and page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector, delay or network idle, blocking ads, trackers, requests or resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, image resizing, selectable cache TTLs, signed links for public image tags, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with parameter names used by other screenshot APIs. These controls help document page state; they do not override a site’s terms or access controls.

Plan Included screenshots per month Price
Free 1,000 $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free, and every feature is available on every plan. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; the MCP server lets AI agents take screenshots; 1,000 screenshots a month are free with no card; and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Validation, freshness, and safe hand-off

Validate the page and the organization

  • Confirm that the canonical domain belongs to the intended organization, especially after redirects or acquisitions.
  • Require a source URL and timestamp for every field used in segmentation or outreach.
  • Check that phone numbers, country codes, industry labels, and postal addresses match your normalization rules.
  • Flag boilerplate, placeholder text, parked domains, and pages that return a successful status but contain an error or consent wall.
  • Keep conflicting observations instead of overwriting one with an unverified “latest” value.

Set a refresh policy

Refresh more often for fast-changing fields such as product availability and less often for stable identity fields. The correct interval depends on the source and purpose; the available sources establish no universal schedule. Expire records that pass their retention deadline and honor deletion or objection requests through the same pipeline that created the data.

Separate enrichment from outreach

Route validated company records to the CRM only after a privacy and marketing review. Keep suppression lists authoritative, record the lawful basis and notice used for a campaign, and make opt-out handling independent of whether a source page remains online.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability, and cost controls

  • Bound the job: cap domains, pages, response bytes, concurrency, and total runtime.
  • Cache carefully: retain a response only for the period your purpose permits, and never use caching to evade a source’s rate limits.
  • Retry narrowly: retry transient network failures with backoff; do not repeatedly retry 403, CAPTCHA, login, or explicit denial responses.
  • Measure operational quality: log status codes, parse failures, redirect chains, exclusion reasons, and review outcomes rather than inventing an accuracy percentage.
  • Control spend: APIs and owner feeds may be more stable than page parsing; managed tools can add subscription and storage costs. Compare total engineering, review, and compliance effort—not just request price.

Troubleshooting common failures

Symptom Likely cause Fix
403, 429, or repeated denials Access policy, rate limit, or an objection. Stop retries, review terms and robots.txt, lower volume only if permitted, and request an API or owner feed.
CAPTCHA or bot-check page The source requires a human or blocks automation. Do not bypass it. Exclude the source or obtain authorized access.
Blank or incomplete fields Client-side rendering, consent overlay, or an error template. Use an authorized API, inspect the rendered state, or capture the page with a rendering service; mark missing values unknown.
Duplicate companies Aliases, country subdomains, redirects, or acquisitions. Normalize domains and names, then send ambiguous matches to review.
Stale enrichment Refresh interval does not match how quickly the source changes. Track timestamps and field-level freshness; schedule a source-specific refresh.
Parser breaks after a redesign Selectors or markup changed. Version the parser, monitor parse-error rates, and prefer structured APIs or stable JSON-LD where authorized.

FAQ

Does a disallow rule prove that scraping is illegal?

No. It is a crawler instruction and an important objection signal, while legality also depends on jurisdiction, terms, rights, data type, and intended use.

Should I store every email address displayed on a company site?

No. Store only fields required for the defined purpose, assess whether an address identifies a person, and apply the privacy, retention, notice, and outreach rules for your recipients.

When is an API worth pursuing?

An API or owner feed is especially valuable when collection is recurring, the data is business-critical, or page changes would create costly maintenance and review work.

Can a screenshot prove that a lead record is accurate?

No. A screenshot documents what a page rendered at a particular time. Accuracy still requires source attribution, normalization, validation, and a freshness policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.