October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI copyright

Web Data for AI and Machine Learning: Where Training Data Really Comes From

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI training data comes from many places—not one master collection of websites. A model’s developers may combine public web crawls, licensed material, public-domain works, human-created examples, platform or user data, and synthetic data. They then filter, clean, deduplicate, and mix those inputs. As a result, a named dataset is usually one stage in a longer pipeline, not proof that every item has the same origin, license, or permission status.

Where does AI training data come from?

Training data can include text, images, audio, video, and other material. For a text model, sources may include web pages, books or other licensed collections, public-domain works, and demonstrations written or rated by people. For image and multimodal models, a training set may pair images with text or include other media. Developers may also generate synthetic examples, then use them alongside other data.

The precise mix depends on the model and its purpose. OpenAI’s public explanations, for example, describe a combination of publicly available information, licensed data, human-created training data, and synthetic data across several modalities. Apple describes directly licensed material, public-domain data, and material available under licenses that permit AI development. Those descriptions are useful categories, not exhaustive inventories of every record used to train every model version.

The pipeline matters as much as the initial source. A builder might collect a crawl, extract text, identify languages, filter low-quality or unwanted material, remove duplicates, and combine the result with other datasets. Each transformation can change what remains. A dataset name therefore does not, by itself, tell you exactly which page, image, or version was included—or whether that item was removed later.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Is ChatGPT trained on web pages?

OpenAI’s public descriptions include publicly available information among the sources used to develop its models, so web information is part of the broad picture. But those descriptions do not provide a complete public, page-by-page list of all URLs used for every ChatGPT model or version. They describe source categories and processing controls rather than an exhaustive inventory of pages, dates, and filtering decisions.

OpenAI also explains that a model consists of parameters (weights) and code that uses them. A trained model is not simply a searchable copy of every page in its training material. That does not establish that a model can never reproduce memorized material; it means that knowing a model was trained on web data is not the same as having a list of its training URLs or being able to retrieve a source record from the model itself.

OpenAI says it filters and processes training material and describes use of robots.txt controls by website owners. Apple likewise says Applebot respects standard robots.txt directives that publishers can use to direct it not to crawl a site or not to use its content to train foundation models. These are company-described operational policies; they do not establish that every organization follows the same rules or that a robots.txt instruction settles every legal question.

What are Common Crawl and C4?

Common Crawl is a source archive

Common Crawl describes itself as a free, open repository of web crawl data. Its data is hosted through Amazon Web Services public datasets, including the s3://commoncrawl/ bucket in us-east-1. Researchers and dataset builders can use crawled web material as an input, but a crawl archive is not the same thing as a finished training dataset: downstream users choose what to extract, retain, filter, or combine.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In a 2024 UK copyright and AI consultation submission, Common Crawl estimated that its archive was a source of 70–90% of tokens used in training data for nearly all of the world’s large language models. That is Common Crawl’s estimate, not a universally verified measurement of every model’s training data. It is best read as an indication of the archive’s claimed influence, not proof that a particular model used a particular page.

C4 is a filtered derivative

C4, or the Colossal Cleaned Crawled Corpus, is a text corpus made from a Common Crawl snapshot. It illustrates how a broad crawl can become a named dataset after processing. Research documenting C4 found text from sources that might surprise readers, including patents and U.S. military websites. A 2025 Creative Commons analysis reports that C4 content originated from more than 14 million web domains. That breadth can encompass reference pages, forums, news sites, businesses, personal pages, government material, and many other kinds of web content.

Those findings also show why “cleaned” should not be mistaken for “licensed,” “representative,” or “safe for every use.” Cleaning describes a processing step; it does not guarantee that all sources have a common license, that the resulting dataset contains no personal information, or that every use is lawful in every jurisdiction.

How do web images enter AI training?

Image datasets may connect an image to text found with it on a web page. LAION-400M documents 400 million English image-text pairs, extracted from Common Crawl pages crawled between 2014 and 2021. LAION provides metadata and links; users must redownload the images themselves. Licensing information may be incomplete or uncertain for an individual image.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LAION’s 2023 maintenance note describes LAION-5B as containing more than 5.85 billion entries. It says the dataset is sourced from the Common Crawl index and provides links to public-web content rather than hosting image files. That distinction matters: an index that points to a file, the site hosting that file, and a model developer that may later use it are separate stages with potentially different records and responsibilities. A link in a dataset is not itself proof of the image’s license or of a model’s use of it.

Are C4 and LAION copyrighted?

There is no reliable one-word answer for every item in either dataset. A dataset can contain or point to material from many creators and sites, with different rights and terms. Public availability is not the same as permission for every downstream use, and a dataset label does not establish that every item is licensed for commercial training.

For a specific item, examine the original work and its host, the applicable license and terms of service, the jurisdiction, and the dataset builder’s stated collection and removal practices. Consider whether personal data is exposed, how robots.txt signals were treated, and whether a relevant text-and-data-mining exception applies. Copyright and text-and-data-mining rules vary by country; an organization may also choose policies stricter than the minimum it believes the law requires. Dataset provenance and legal conclusions are related, but they are not interchangeable.

Can you find the exact websites used to train a model?

Sometimes you can trace a dataset or a sample of its contents to source websites; that is different from identifying every URL used by a proprietary model. Public disclosures from OpenAI and Apple describe source categories and controls, but they do not publish a complete page-level list for every model version, URL, or filtering threshold. A model developer’s use of a public corpus also does not mean the corpus’s entire contents were retained for training.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For open or documented datasets, check whether the release preserves source URLs, crawl dates, derivation records, and version information. Even then, links can stop working, source pages can change, and the existence of a URL in an index does not prove that a particular model consumed that record. Treat claims about exact model training pages cautiously unless they are tied to records for the relevant dataset, model, and version.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to check a dataset’s provenance and license

Use the same questions whether you are assessing a Common Crawl derivative, a LAION release, a licensed collection, or a proprietary mixture. Record what is documented and what remains unknown rather than treating “open” or “public” as a complete rights assessment.

  1. Identify the exact release. Record the dataset name, version, publisher, release date, and, where available, a hash or other identifier. A name without a version may refer to different contents or documentation.
  2. Trace origin and lineage. Look for original URLs or record identifiers, collection dates, source corpora, and each documented transformation. Ask whether derivation links survive filtering and whether the release distinguishes original source data from later annotations.
  3. Understand the contents and scale. Check modality (text, image-text, audio, video), item or token counts, and language coverage. Find out whether counts refer to URLs, files, pairs, entries, or tokens; these units are not interchangeable.
  4. Read the license and terms closely. Determine what license applies to the dataset itself and whether it applies to underlying works. Look for direct licenses, public-domain claims, open licenses, restrictions, and any stated limits on commercial use. If an item’s rights are uncertain, do not infer permission from its presence in the corpus.
  5. Inspect collection and filtering rules. Look for language identification, quality filters, safety classifiers, deduplication, personal-data handling, robots.txt policy, and publisher opt-out or removal processes. Note what is not explained; a missing detail is not evidence that a control was applied.
  6. Check reproducibility and correction paths. Prefer versioned releases with documentation, code, hashes, and a way to report errors or request correction or removal. Establish who receives a request and whether it can affect an already released dataset or only future releases.
  7. Assess freshness and drift. Compare the collection period with the date you need to represent. Websites change, disappear, or change ownership and licensing terms; a captured crawl is a record of a past state, not a guarantee about the current page.

The Data Provenance Initiative’s Explorer offers a useful model for the kinds of records to seek: its project description says it tracks sources, licenses, creators, geographies, modalities, and derivation chains across more than 4,000 datasets. That index can help locate documentation, but the underlying dataset and source records still need to be checked for the question you are asking.

Preserving a source page during a provenance review

A screenshot can help a reviewer retain a visual record of what a source page looked like at capture time. It cannot establish the page’s license, prove that it appeared in a training set, or replace a URL, crawl timestamp, dataset version, and rights review. Keep those records alongside any image you save.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

For a visual record, ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API can return a PNG, JPEG, WebP, or PDF; cookie banners, popups, and chat widgets are removed before capture, and each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, with the response identifying the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDFs. It includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. See ScreenshotNeo and the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Replace the example target URL with the page you are reviewing, and keep the capture associated with that URL and your own date and dataset records. Sign up for 1,000 free screenshots a month with no card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.