Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The best starting point depends on your experiment. Use Phishpedia when you need a published visual benchmark with URL, HTML, screenshot, and target-brand annotations; use the Zenodo “Phishing and Legitimate Websites Dataset” when you need a mixed phishing/legitimate collection with PNG screenshots and CSV features; use PhishTank or OpenPhish to discover current suspicious URLs, not as guaranteed, fixed screenshot benchmarks. PhishVN is useful for time-stamped Vietnamese data when its access and handling conditions fit your work.

A screenshot is evidence of what a browser rendered at capture time—not proof that the site was still online when you analyze it. Build capture timestamps, liveness results, labels, provenance, and safe-isolation procedures into your dataset from the beginning.

Which phishing screenshot dataset should you choose?

Resource What is documented Best fit Important qualification
Phishpedia benchmark About 30,000 phishing webpages, each described with a URL, HTML, screenshot, and target brand. Visual phishing identification, brand-impersonation and multimodal research. Confirm the current repository release, labels, download access and reuse terms before building a benchmark.
PhishTank Verified or online phishing URL data; detail records may show screenshots and community votes. Finding candidate URLs, URL-feed integration and building a capture pipeline. A feed is not a fixed screenshot corpus. Screenshot availability and capture state must be checked per record.
OpenPhish Database Structured phishing indicators, infrastructure information, tiered update cadence and retention options. Current threat-intelligence or URL/host analysis, including training or validation workflows where your access tier permits it. The documented fields are indicators, not a guaranteed collection of webpage images. Access and pricing vary by tier.
Phishing and Legitimate Websites Dataset (Zenodo) A July 15, 2026 record describing 60,000 URLs: 31,641 phishing and 28,359 legitimate, with PNG screenshots and CSV features. Mixed-class experiments using image and tabular inputs. Verify the exact version, files, license and capture methodology before quoting the counts as your own corpus.
PhishVN A time-stamped Vietnamese URL dataset with open and gated tiers; the gated evidence bundle includes rendered HTML and screenshots. The article reports 868 gated records: 209 phishing and 659 benign. Region- and scenario-specific research where timestamps and evidence tiers matter. The gated archive is not openly downloadable; the article specifies research-only handling and isolated-VM precautions for HTML.

What each source can—and cannot—answer

Phishpedia: a visual benchmark with context

Phishpedia’s project description says its roughly 30,000 phishing examples are annotated with the URL, HTML, screenshot and target brand. That pairing lets you study both visual similarity and page context instead of training on pixels alone. The associated USENIX Security 2021 paper provides the research setting for the release. Treat “about 30k” as the project’s description, then record the repository version, download date, label definitions and any filtering you apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PhishTank: candidate discovery, not a frozen image set

PhishTank documents verified or online phishing URLs and detail pages that can include screenshots and community votes. That is useful for selecting URLs to capture, checking report metadata or creating a continuously refreshed pipeline. It does not, by itself, establish that every feed record has an image, that images were captured at the same time, or that the resulting set is versioned and reproducible.

OpenPhish: indicators and freshness

OpenPhish describes a searchable database of phishing indicators with tiered update cadence and retention. This makes it relevant when your question is “which hosts or URLs are active now?” rather than “what does every page in a fixed image benchmark look like?” Document your subscription tier, retrieval time and retention window. Do not silently call an indicator feed a screenshot dataset.

Zenodo’s mixed collection

The Zenodo record published July 15, 2026 states that it contains 60,000 website URLs, split into 31,641 phishing and 28,359 legitimate URLs, plus PNG screenshots and CSV features. Those are the record’s stated figures. Before reporting model results, inspect the release files, version, license, class-generation rules, screenshot dimensions, missing values and capture dates.

PhishVN: geography and evidence tiers

PhishVN is appropriate when Vietnamese coverage and timestamps are central to the question. Its open tier and gated evidence bundle are different products: the article describes rendered DOM/HTML and screenshots for the gated records and reports research-only handling requirements. Do not redistribute gated HTML or images without checking the release terms.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to evaluate a screenshot dataset before downloading

Use a written checklist so a visually attractive collection does not hide gaps that invalidate your experiment.

  • Visual evidence: Is an image present for every record, only some records, or only on separate detail pages? Is it a viewport capture, a full-page image, or unspecified?
  • Labels: Does the record distinguish phishing from legitimate pages? Does it identify the impersonated brand, confidence, scenario or verification state?
  • Pairing: Can you reliably join URL, screenshot, HTML, redirects, timestamp and label with a stable record ID?
  • Coverage: What countries, languages, brands, hosting providers and collection period are represented?
  • Freshness: Is this a fixed snapshot or an updating feed? Can you reproduce the same version later?
  • Evaluation design: Are duplicates, near-duplicates, inactive pages, brand overlap and time leakage addressed?
  • Reuse: Is download open, gated, rate-limited or paid? May you redistribute images, HTML, derived features and model weights?

Design a defensible train/validation/test split

Prevent duplicate leakage

Hash the original image, then use perceptual hashes or image embeddings to find resized, recompressed and template-level duplicates. Also normalize URLs and compare registrable domains. A page copied across many subdomains can make a random split look far better than deployment performance.

Split by time when the task is operational

If you want to detect future campaigns, train on earlier captures and test on later captures. Keep the capture timestamp separate from the URL-report timestamp: a screenshot can survive after the phishing site has gone offline.

Control brand and campaign overlap

For brand-target research, hold out brands or campaigns rather than allowing near-identical login templates in every split. Report whether the test set contains brands seen during training.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep provenance columns

At minimum, store a stable record ID, original URL, normalized URL, final URL after redirects, capture timestamp in UTC, source name and version, label, target brand if known, image path, image dimensions, HTTP status, liveness result, redirect chain, and license or access tier. Add a field for “screenshot present” instead of treating a missing file as a negative example.

Capture your own screenshots from URL feeds

A feed-to-image pipeline is useful when you need current examples or when a source supplies URLs but not a complete image archive. Run it in an isolated research environment: phishing HTML and linked resources can be hostile, and the PhishVN article specifically describes isolated-VM precautions for gated HTML.

  1. Ingest and freeze the input. Save the feed response, retrieval time, source version and terms. Never overwrite the original URL list.
  2. Normalize safely. Preserve the original string, create a normalized comparison key, and record redirects rather than discarding them.
  3. Capture with limits. Use a disposable browser or capture service, a bounded timeout, blocked downloads where appropriate, and no analyst credentials. Record viewport, user agent, timezone and JavaScript settings.
  4. Record page state. Save screenshot, final URL, HTTP status, redirect chain, load outcome, capture timestamp and whether a consent wall, bot check or blank page appeared.
  5. Quarantine artifacts. Store HTML and assets away from normal workstations. Do not open downloaded files outside the isolated environment.
  6. Deduplicate and split. Apply URL, image and template deduplication before creating train, validation and test sets.
  7. Publish a manifest. Include checksums, schema, source terms, version, filtering rules and a statement that screenshots do not prove contemporaneous liveness.

Minimal Python capture-record pattern

The following pattern illustrates metadata handling; use an isolated browser or a screenshot API for the actual render.

from datetime import datetime, timezone
import hashlib, json

url = "https://example.invalid"
record = {
    "record_id": hashlib.sha256(url.encode()).hexdigest()[:16],
    "original_url": url,
    "captured_at": datetime.now(timezone.utc).isoformat(),
    "source": "your-feed-name",
    "screenshot_present": False,
    "liveness": "unknown"
}
with open("manifest.jsonl", "a", encoding="utf-8") as f:
    f.write(json.dumps(record) + "n")

Common failure modes and fixes

“The dataset says screenshots, but files are missing.”

Check whether images are in a separate archive, require authentication, or are available only on per-URL detail pages. Count missing files explicitly and do not label them legitimate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The screenshot is a blank page or a bot check.”

Record that state as a capture outcome. Do not silently retry until a page appears; repeated retries can bias the sample toward pages that tolerate your browser.

“The page was offline during validation.”

Keep the historical screenshot, but set liveness to unknown or inactive at validation time. A 2021 study of PhishTank screenshots found examples captured after the phishing site had become inactive.

“Model accuracy collapses on a new campaign.”

Inspect template and brand leakage, then evaluate a time-held-out split. Include redirects, language, viewport and capture-period fields in error analysis.

“Redistribution is blocked.”

Separate code and derived statistics from raw HTML and images. Re-read the exact release terms; open access to a URL list does not automatically grant permission to redistribute rendered evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is the first option to try when you need a screenshot API: it removes cookie/consent banners, newsletter popups and chat widgets before capture, bills only clean shots, and does not bill bot checks, CAPTCHAs, blank pages, timeouts, failed loads or cache hits. Responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page and CSS-selector captures, dark mode, device presets, arbitrary viewports, retina scale, PDF paper settings, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.

cURL (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan, and yearly billing provides two months free. Create a free ScreenshotNeo account to begin.

Cost, performance and reproducibility notes

  • Cache deliberately: A cache TTL can reduce repeat captures, but cached output is not a new observation. Store whether a result came from cache.
  • Use asynchronous jobs for volume: Webhooks avoid holding a worker open for slow pages. Persist job IDs and signed callback details in your manifest.
  • Separate throughput from coverage: Parallel requests improve speed but can trigger bot defenses or overload a host. Set bounded concurrency and retries.
  • Measure outcomes, not only files: Count clean pages, bot checks, blank pages, timeouts, failed loads and cache hits separately.
  • Keep the capture recipe: Browser version, viewport, locale, timezone, wait rule, blocked resources and custom scripts can all change pixels.

FAQ

Can I call PhishTank a phishing screenshot dataset?

Only for the specific records for which you have verified screenshots and capture metadata. Its documented feed is primarily URL and verification data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Are screenshots safe to open?

The image itself is less risky than executing page HTML, but treat the surrounding files, links and embedded resources as potentially hostile and use isolation.

Should I combine phishing and legitimate screenshots?

Yes, when your task is binary classification, but document how legitimate pages were sampled and ensure the classes do not differ only by source, geography or capture date.

Frequently Asked Questions

Can I call PhishTank a phishing screenshot dataset?

Only for records where you have verified screenshots and capture metadata; its documented feed is primarily URL and verification data.

Are screenshots safe to open?

Treat associated HTML, assets and links as potentially hostile and use an isolated environment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I combine phishing and legitimate screenshots?

Yes for binary classification, provided you document legitimate-page sampling and control source, geography and time confounders.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.