What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
You can build a useful B2B prospect database without buying a massive contact file: define the account and role fields you actually need, collect company facts only from sources whose terms permit your planned access and reuse, preserve provenance, validate and deduplicate records, then review the marketing rules before contacting anyone. A page being publicly visible is not blanket permission to automate collection or reuse its contents.
Start with a prospect definition, not a scraper
Write the database specification before choosing a library or browser. A clear specification prevents collecting personal details that you cannot justify and makes duplicate detection possible.
Account fields
| Field | Purpose | Collection note |
|---|---|---|
| Legal or trading name | Identifies the organization | Store the spelling shown by the source and a normalized comparison value. |
| Primary website and domain | Deduplication and account routing | Normalize protocol, case and trailing slashes; retain the original URL as provenance. |
| Headquarters or operating region | Territory and regulatory review | Record the source wording and do not infer a person’s location from it. |
| Industry or product category | Segmentation | Use a controlled vocabulary plus the source’s original description. |
| Size signal | Qualification | Label whether it is self-reported, estimated or not stated. |
| Technology, use case or buying signal | Relevance | Capture only an observable business fact and its date. |
Person-level fields
Collect an identifiable employee’s name, job title or direct address only when a stated business purpose requires it. Keep the minimum necessary fields, record how the information was obtained, and create a process for objections, correction and deletion. A company domain is not the same thing as permission to profile every employee at that company.
Check permission before automating access
Review each source’s rules
Read the site’s terms, access restrictions, API conditions and any stated reuse limits before writing a collector. Check whether automated requests are allowed, which paths are excluded, rate limits, attribution requirements and whether storing or redistributing the resulting data is restricted. Treat a login wall, CAPTCHA, technical block or explicit prohibition as a stop signal, not an obstacle to bypass.
#1 Best Overall
CNIL explains that scraping is not inherently incompatible with GDPR requirements, but other rules can prohibit particular activity, including terms based on database-producer rights or copyright law. That means the answer depends on the source, the fields, the purpose, the people involved and the jurisdictions—not merely on whether a browser can see the page.
LinkedIn is not an acceptable shortcut
LinkedIn’s published policy expressly prohibits third-party crawlers, bots, browser extensions and other methods used to scrape or copy its services, including profiles. It warns that accounts can be restricted or shut down. Do not scrape LinkedIn profiles or evade its controls. In a May 6, 2022 company statement about Mantheos, LinkedIn said Mantheos agreed to delete scraped profile data and stop automated access. That is a platform-enforcement example, not a universal legal precedent for every website.
Separate company research from personal-data collection
- Prefer organization-level facts: products, locations, published capabilities, certifications and company contact channels.
- Document why each person-level field is needed for the stated outreach or qualification task.
- Do not guess an email address from a naming pattern and present it as verified.
- Keep source URL, collection date, fields collected, purpose and the permission or legal basis assessed with every record.
A defensible collection workflow
- Set inclusion rules. Define industries, regions, account size signals and disqualifiers. Decide which fields are mandatory and which are optional.
- Make a source register. For every domain, record the terms reviewed, access method, allowed frequency, reuse limits, owner and review date.
- Choose the least intrusive method. Use a permitted API or export when available. Static HTML requests are simpler and cheaper than a browser. Use browser automation only when a permitted page genuinely requires client-side rendering.
- Collect with provenance. Save the exact source URL, timestamp, parser version, fields extracted and a hash or archived copy where your retention policy allows it.
- Validate. Check required fields, normalize domains, verify that the page still describes the same organization and flag contradictory values for human review.
- Deduplicate. Match first on a normalized domain, then on a normalized legal name plus country or region. Never merge two accounts solely because their names are similar.
- Review retention and outreach. Set a review cadence based on how quickly the data changes, honor correction or deletion requests, and assess the recipient’s and sender’s jurisdictions before sending.
DIY extraction with a permitted public page
The following example collects a page title, description and heading for an organization-level record. Replace TARGET_URL only with a URL whose terms permit your access and reuse. It intentionally does not harvest personal profiles or guessed email addresses.
Python: requests and Beautiful Soup
import os
from datetime import datetime, timezone
from urllib.parse import urlparse
import requests
from bs4 import BeautifulSoup
url = os.environ["TARGET_URL"]
headers = {"User-Agent": "B2BResearchBot/1.0 (contact: [email protected])"}
response = requests.get(url, headers=headers, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else ""
meta = soup.find("meta", attrs={"name": "description"})
description = meta.get("content", "").strip() if meta else ""
h1 = soup.find("h1")
record = {
"source_url": url,
"source_domain": urlparse(url).netloc.lower(),
"collected_at": datetime.now(timezone.utc).isoformat(),
"page_title": title,
"description": description,
"primary_heading": h1.get_text(" ", strip=True) if h1 else "",
}
print(record)
Install dependencies with python -m pip install requests beautifulsoup4. Add a queue, a delay and bounded retries rather than sending a burst of requests. Persist the response status and failure reason so a timeout is not mistaken for a missing company.
cURL: save the permitted HTML for inspection
export TARGET_URL="https://example.com/allowed-page"
curl -L --max-time 30 -A "B2BResearchBot/1.0" "$TARGET_URL" -o page.html
Node.js: fetch and record basic metadata
const target = process.env.TARGET_URL;
if (!target) throw new Error('Set TARGET_URL');
const response = await fetch(target, {
headers: { 'User-Agent': 'B2BResearchBot/1.0' },
signal: AbortSignal.timeout(30000)
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const title = html.match(/<title[^>]*>([sS]*?)</title>/i)?.[1]?.trim() ?? '';
console.log(JSON.stringify({ source_url: target, status: response.status, title, bytes: html.length }));
These examples are collectors, not permission. Add robots and terms checks, concurrency limits, and a manual review path before running them against a list of domains.
When a browser is necessary
Some permitted sites render the relevant company facts only after JavaScript runs. A browser can wait for a selector, click a tab or choose a locale, but it also increases cost, failure modes and the amount of data that may be exposed to your process. Use a fixed viewport, a short maximum wait, blocked nonessential resources and a low concurrency. Never use automation to defeat a CAPTCHA, bot check, paywall or access restriction.
Capture the final URL, response status, console errors and a screenshot or HTML artifact when your governance policy allows it. A failed load should become a review status—not an empty prospect record.
Validate, score and maintain the database
Quality checks
- Required account name and source URL are present.
- Domain resolves and is not a parking, redirect-only or disposable domain.
- Industry and region values match the source text.
- Conflicting records are queued for review instead of silently overwritten.
- Every value has a collection timestamp and source.
Deduplication keys
Normalize Unicode, case, punctuation and common legal suffixes for comparison, while retaining the original display value. Use the registrable domain as the strongest key when it represents one business. For groups, franchises and subsidiaries, preserve separate entities when their buying authority or geography differs.
Recommended Free Tools
Refresh and deletion handling
There is no universal retention period established for every geography or prospect type. Set one appropriate to your purpose and applicable rules. Recheck fast-changing fields more often than stable facts, mark records as stale rather than silently changing history, and record when an objection or deletion request was received and completed.
Rules for B2B cold email
In the United States, the FTC says CAN-SPAM applies to commercial messages, including B2B email. Its business guide requires accurate header information, non-deceptive subject lines, identification that the message is an advertisement, a valid physical postal address and a working opt-out method. The guide states: “That means all email – for example, an email promoting a product or service to former customers – must comply with the CAN-SPAM Act.”
Rank #3
Before sending, separately assess the sender’s and recipient’s jurisdictions, the channel and the nature of the data. The available guidance does not establish one worldwide legal basis, notice rule or retention period. Regional privacy and marketing review may be necessary, especially for person-level data.
What if a vendor sends the email?
Outsourcing delivery does not transfer the business’s compliance responsibility. The FTC guide says a company cannot contract away that responsibility. Give the vendor accurate suppression lists, monitor opt-outs and retain evidence that your messages used truthful sender details and a valid postal address.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Reliability, performance and cost decisions
| Method | Strength | Trade-off |
|---|---|---|
| Manual review | Best for ambiguous sources and small lists | Slow and difficult to repeat consistently |
| Permitted API or export | Stable fields and documented limits | May omit fields or impose quotas and reuse restrictions |
| Static HTTP requests | Fast, inexpensive and easy to log | Cannot see facts rendered only in JavaScript |
| Browser automation | Handles permitted interactive pages | Higher runtime, more breakage and greater governance burden |
Use caching for unchanged pages where the source permits it, exponential backoff for transient failures, bounded concurrency and idempotent jobs. Track success rate, status codes, parse failures, average response time and records requiring review. A cheaper request that produces stale or merged companies is not a lower-cost database.
Or skip the browser setup: ScreenshotNeo for visual source records
If your workflow needs a visual record of a permitted company page rather than a custom browser stack, ScreenshotNeo is the first screenshot API to try: it removes cookie banners, newsletter popups and chat widgets before capture, bills only clean shots, and has the lowest paid plan listed here.
One GET request returns a PNG, JPEG, WebP or PDF. The API can capture full pages with lazy images loaded, one CSS-selected element, dark mode, 12 device presets or any viewport, retina scale, paper-size PDFs with margins, custom CSS and JavaScript, a click before capture, selector or network-idle waits, blocked ads/trackers/requests/resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options and response headers. Each response identifies the page verdict and whether it was billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #4
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Plans include 1,000 shots per month free with no card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start.
Troubleshooting common failures
HTTP 403, 429 or an account warning
Stop the job, reread the source rules and reduce frequency. Do not rotate identities, bypass controls or continue against an explicit prohibition.
The page is blank or missing fields
Check whether content is JavaScript-rendered, whether a consent state is required, and whether the request received a redirect or an access challenge. Use a permitted API, a documented export or a human review instead of trying to defeat the challenge.
Many duplicate companies
Normalize domains and legal names before insertion, preserve subsidiaries separately, and require a reviewer for fuzzy matches.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRecords become stale
Store collection dates and parser versions, schedule refreshes according to field volatility, and mark unverified records for recheck.
Best Value
Opt-outs keep reappearing
Maintain a central suppression list keyed by the address and, where appropriate, the organization. Apply it before every campaign and synchronize it with any delivery provider.
FAQ
Should I store the original HTML?
Only when your documented purpose, source terms and retention policy allow it. Otherwise retain the minimum structured fields plus provenance and a review trail.
How should I choose a refresh interval?
Base it on how quickly each field changes and the risk of acting on stale information; there is no universal interval that fits every market or jurisdiction.
Can a screenshot replace provenance?
No. A screenshot can show what a permitted page looked like at a time, but your record still needs the source URL, collection date, fields used and the purpose for collecting them.
Frequently Asked Questions
Should I store the original HTML?
Only when your documented purpose, source terms and retention policy allow it. Otherwise retain the minimum structured fields plus provenance and a review trail.
How should I choose a refresh interval?
Base it on how quickly each field changes and the risk of acting on stale information; there is no universal interval that fits every market or jurisdiction.
Can a screenshot replace provenance?
No. A screenshot can show what a permitted page looked like at a time, but your record still needs the source URL, collection date, fields used and the purpose for collecting them.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




