October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI

How AI Companies Collect Training Data with Web Crawlers—and What Publishers Can Control

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI training collection is a pipeline, not a single download. A crawler discovers URLs, requests pages, records response and provenance data, extracts and normalizes content, then applies permission, quality, privacy and duplication filters before material enters a dataset. A page being publicly reachable does not by itself grant unrestricted permission to copy or reuse it. Publishers should treat robots.txt as an operational instruction, document separate search and training choices, review contractual and copyright obligations, and keep auditable records of crawler access.

The web-crawling pipeline used for training data

Implementations differ by company, but a modern collection system normally has these stages:

  1. URL discovery. A crawler seeds a URL frontier from previously known pages, links found in fetched documents, sitemaps, feeds, public submissions and other permitted sources. The frontier stores URLs that still need a request and applies scheduling and rate limits.
  2. Permission and request checks. Before fetching, the crawler identifies its user-agent, checks robots.txt rules and applies internal allow, deny or opt-out policies. Different products may use different bots for search, model training, advertising or a user-triggered fetch.
  3. Fetching. The crawler requests the page and records status code, final URL after redirects, response headers, content type, retrieval time and errors. It may retain the raw response, subject to its policy and legal basis.
  4. Parsing and normalization. HTML is converted into text and structured fields. Boilerplate, navigation, duplicate markup, scripts and other non-content elements are identified. Links, language, publication dates and page metadata can be extracted for later processing.
  5. Filtering and safety review. Systems remove or down-rank spam, malware, low-quality duplicates and categories of unwanted personal data. OpenAI describes using publicly available webpages, forums, blogs and posts while applying filters that include spam and some unwanted personal-data sources; that description is not a universal recipe for every provider.
  6. Dataset assembly and provenance. Accepted records are deduplicated, assigned source and retrieval metadata, and stored with policy decisions or lineage information so later training or evaluation can be traced to an input and collection date.

These stages are often repeated: pages are recrawled, policy files are rechecked, and records can be removed when a source changes its terms or submits a valid request.

What “publicly available” does—and does not—mean

A page that loads without authentication is technically accessible to a crawler, but accessibility is not the same as a blanket license. A collection decision can implicate:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Copyright and related rights. The U.S. Copyright Office’s AI initiative is examining copyright questions raised by training on protected works. Its report is being issued in parts, including a generative-AI-training part released in 2025. Outcomes remain dependent on the jurisdiction, facts, purpose and parties involved.
  • Contract terms. Terms of service, API agreements, paywalls, license notices and database-rights rules may impose conditions that robots.txt cannot override.
  • Privacy and data protection. Public exposure does not eliminate duties concerning personal information. Collection programs need retention, minimization, access, deletion and security controls appropriate to the jurisdictions and data involved.
  • Publisher instructions. Robots.txt, meta directives, headers and provider-specific opt-out channels can communicate operational preferences. They should be recorded with the other legal and governance signals rather than treated as the only permission record.

Because these questions are fact-specific, a publisher should obtain advice for its jurisdictions and content model instead of relying on a universal claim that web training is either always lawful or always prohibited.

How robots.txt controls crawler behavior

Google documents that its crawlers download and parse robots.txt before crawling and select the most specific matching user-agent group. In practice, a crawler compares its identity with the groups in the file, then applies the matching Allow and Disallow rules. A syntax error, an unreachable file or a rule that does not match the bot’s exact identity can produce a different result from the one a publisher intended.

Robots.txt is an operational signal, not a copyright license, contract, consent record or legal waiver. A crawler may be technically able to fetch a URL despite a rule, and a rule alone does not settle whether copying or later use is lawful. Keep a dated copy of every policy change and the reason for it.

Propagation is not instantaneous. OpenAI says robots.txt changes can take about 24 hours to affect its search crawling behavior. Allow time for caching and recrawling, and verify requests in your server logs after the expected interval.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Separate search visibility from model-training access

OpenAI publishes separate controls for OAI-SearchBot and GPTBot. GPTBot is associated with content that may be used to train foundation models, while OAI-SearchBot is used for search presentation. OpenAI’s documentation emphasizes: “Each setting is independent of the others.” That means a publisher can choose to allow search crawling while disallowing the training-associated bot, subject to the limits of robots.txt enforcement and any other applicable policies.

A conceptual policy might look like this:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Use the exact user-agent names and syntax documented by the provider, test the result against your production robots.txt, and confirm that your CDN or WAF does not rewrite or cache an older file. This example communicates intent; it does not create a legal permission or guarantee that every intermediary will honor it.

What Common Crawl provides, and the legal limit of its terms

Common Crawl describes a corpus with three principal layers: raw web-page data, metadata extracts and text extracts. Those layers can support different workflows, from reprocessing original responses to searching text and studying crawl metadata.

Its terms permit use in connection with AI systems, including developing, training or deploying them. The same terms warn that crawled material can carry separate terms and third-party rights and require compliance with applicable law. Using a Common Crawl file therefore does not transfer ownership of every page in it, erase a publisher’s license conditions or resolve privacy obligations. A model developer still needs controls for source restrictions, removal requests, personal data and downstream distribution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to compare training-data collection approaches

No single source or corpus wins on every dimension. Evaluate the collection system and the resulting dataset against the same questions:

Evaluation axis Questions to ask
Permission and opt-out handling Does the collector identify bots clearly, parse robots.txt correctly, honor provider-specific opt-outs and retain evidence of each decision?
Coverage Which domains, languages, regions, content types and accessibility levels are represented, and which are systematically absent?
Freshness How often are pages recrawled, how are changed or deleted pages detected, and can the dataset distinguish retrieval dates?
Filtering and deduplication What spam, malware, boilerplate, near-duplicate and low-quality filters are applied, and can false positives be investigated?
Personal-data minimization Are sensitive fields detected, removed or masked before release and training? What retention and deletion process exists?
Provenance and reproducibility Can a record be tied to a URL, retrieval time, response metadata, transformation and policy decision without exposing unnecessary personal information?
Licensing and downstream use What rights attach to the source, the extracted text, the trained model and any redistributed dataset?
Infrastructure and rate limits How are concurrency, retries, politeness delays, bandwidth, robots failures and provider blocks handled?

Documenting these answers is more informative than comparing a single corpus-size number, especially because no authoritative, universal corpus-size figure is established for the sources described here.

A publisher’s practical control and audit workflow

  1. Inventory crawler identities. List observed user-agent strings and classify them as search, training, advertising, monitoring or user-triggered access. Treat an unknown bot as unknown until verified through published documentation and request logs.
  2. Publish and test robots.txt groups. Create separate groups for the bots you can identify, test matching and precedence, and keep a dated change log. If search and training choices differ, express them in separate groups rather than a single broad rule.
  3. Review terms and licenses. Add terms-of-service, syndication licenses, API agreements, privacy notices and opt-out commitments to the crawl-approval checklist. Robots rules should be one recorded signal among these documents.
  4. Log every material decision. Retain request time, user-agent, URL, response status, redirects, robots result, policy version and opt-out evidence. Limit retained content and personal data to what the governance purpose requires.
  5. Filter before release. Apply spam and malware screening, deduplication, quality checks and personal-data minimization before a dataset is shared or used for training. Record filter versions so a later audit can reproduce the decision.
  6. Recheck continuously. Revisit policies when crawler behavior, standards interpretation, provider documentation or copyright rules change. A 2024 NeurIPS Datasets and Benchmarks study tracked robots.txt and terms-of-service restrictions for major AI developers and web archives from 2016 through April 2024, illustrating why longitudinal audits are useful.

Capture visual evidence of what a page presented

For an access audit, a publisher may want a visual record of a page before and after a consent banner, newsletter prompt or chat widget appears. A do-it-yourself browser automation script can save screenshots at defined checkpoints, while server logs remain the authoritative record of requests and bot identities. Keep screenshots as evidence of presentation, not as proof that a crawler had legal permission to copy the underlying work.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. It can accept a page’s cookie or consent banner like a visitor and remove more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the one-call endpoint for a visual checkpoint (replace the target URL with the page you are auditing):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for parameters and response headers. Equivalent Python:

Rank #4
Web-Crawler
  • SUPERHERO AND VEHICLE FIGURE SET: Many adventures with this Spidey and His Amazing Friends set, which includes a figure, vehicle, and accessory
  • ARTICULATED FIGURE: This 4" figure features multiple points of articulation for lots of action
  • TEAM SPIDEY ADVENTURES: Kids can be part of Team Spidey and create their own epic adventures with this Spidey and His Amazing Friends Vehicle Set
  • INSPIRED BY MARVEL'S CHILDREN'S DRAWING: Little kids can imagine saving the day with their favorite superheroes with this Spidey and His Amazing Friends toy, inspired by the cute kids show
  • ENDLESS ADVENTURES WITH SPIDEY AND HIS AMAZING FRIENDS TOYS: Other Spidey and His Amazing Friends Toys Available (sold separately and subject to availability)
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For audit scenarios, options include full-page captures with lazy images loaded, a CSS-selected element, dark mode, device and retina settings, custom CSS or JavaScript, selector hiding, waits for a selector, delay or network idle, blocked ads or trackers, custom headers and cookies, timezone and geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which can simplify migration.

ScreenshotNeo’s MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Plans are: Free, 1,000 shots per month with no card; Starter, $5 for 3,000; Growth, $15 for 15,000; Pro, $39 for 60,000; Scale, $99 for 250,000; and Business, $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Start with 1,000 free screenshots a month—no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failure modes and fixes

The wrong bot is being allowed or blocked

Cause: a user-agent group is misspelled, too broad or overridden by a more specific group. Fix: compare the exact logged user-agent with the documented token, test precedence and publish a corrected file.

Changes appear to have no effect

Cause: cached robots.txt content or a crawler that has not recrawled yet. Fix: verify the file at the production origin and CDN, retain the change timestamp, allow the documented propagation window and inspect subsequent requests.

A dataset still contains a page that opted out

Cause: the page was collected before the opt-out, entered through another source such as an archive, or was not linked to its provenance record. Fix: trace the URL and retrieval date, apply the removal process to every derivative, and record the decision so future training jobs exclude it.

Logs show a crawler but no reliable permission record

Cause: access logging was separated from policy and terms review. Fix: join request logs to the robots.txt version, terms snapshot, opt-out evidence and dataset record; quarantine material that cannot be evaluated until governance review is complete.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Bottom line

AI companies generally discover, fetch, normalize, filter and provenance-tag web content before assembling training data. Publishers can influence that pipeline with precise bot policies, contractual and licensing controls, privacy safeguards and auditable logs. robots.txt helps communicate operational intent, but it is neither a complete legal answer nor a substitute for governance. OpenAI’s separate OAI-SearchBot and GPTBot controls let publishers make search and training choices independently, while Common Crawl’s AI-use terms still leave responsibility for third-party rights and applicable law with the user.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.