Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Short answer: OpenAI did not literally admit that it had a legal right to use every copyrighted work for free. In a 2024 submission reported by Futurism, it argued that training leading, general-purpose AI models solely on public-domain material would not meet contemporary users’ needs, and that copyright law does not automatically prohibit AI training. That is a policy and legal position—not a final ruling that all training copies are lawful.
What OpenAI actually argued
The headline “OpenAI pleads that it can’t make money without using copyrighted materials for free” is an inflammatory paraphrase of a narrower claim. OpenAI told a House of Lords Communications and Digital Committee that “it would be impossible to train today’s leading AI models without using copyrighted materials.” It also argued that public-domain books and drawings alone would not provide the breadth and quality needed by modern users.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Copyright Law | $145.13 | Buy on Amazon |
| 2 |
|
Copyright Law: Cases and Materials (v8.0) | $21.70 | Buy on Amazon |
| 3 |
|
Copyright Law of the United States: and Related Laws Contained in Title 17 of the United States Code | $10.32 | Buy on Amazon |
| 4 |
|
Copyright Law in a Nutshell | $65.00 | Buy on Amazon |
| 5 |
|
Copyright Handbook, The: What Every Writer Needs to Know | $37.99 | Buy on Amazon |
The important qualifier is leading, contemporary models. OpenAI was not claiming that no AI system could be built from public-domain or openly licensed material. Its argument was that such a restriction would make it difficult to build models with the capabilities, coverage and commercial competitiveness expected of general-purpose systems.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteOpenAI also maintained that copyright law does not categorically forbid training. It distinguished training from distributing a verbatim copy of a book, article, photograph or song. The company later repeated a related economic argument in written evidence to Parliament: mandatory licensing could be expensive, difficult to administer and more accessible to large incumbents than to startups. See OpenAI’s written evidence.
#1 Best Overall
Why copyrighted material is difficult to avoid
“Copyrighted material” does not mean that every byte on the internet is protected in the same way. A training corpus can contain:
- Public-domain works: material whose copyright has expired or that was never protected.
- Licensed works: content used under negotiated or standardized permissions.
- User-provided material: data supplied under terms that may or may not permit model training.
- Factual and government material: facts, raw measurements and some government works may receive limited or no copyright protection, while the expressive presentation of those facts may be protected.
- Web content: publicly accessible pages can still contain protected articles, photographs, illustrations, software and other expression.
- Synthetic data: material generated by other models, which can bring quality, provenance and recursive-training problems of its own.
Modern models need large and varied collections of language, images, code and other human expression. OpenAI’s position is that restricting developers to public-domain material would remove much of the contemporary data that makes models useful. That is a claim about capability and data availability, not proof of legal necessity.
The copyright dispute is about a pipeline, not one event
The legal question is often presented as “does a model contain a copy of a book?” That is only part of the analysis. At least four separate issues can arise:
- Acquisition: Was the work obtained lawfully, or through piracy, unauthorized access or a breach of applicable terms?
- Intermediate copying: Did downloading, storing, preprocessing or tokenizing the work create copies that engage copyright rights?
- Training: Is the use covered by an exception, sufficiently transformative or otherwise lawful?
- Output: Does the system reproduce protected expression, memorize passages or generate material that substitutes for the original market?
A court could find some training activity lawful while still imposing liability for unlawful source acquisition, particular outputs, memorization or failures involving protected works. Conversely, the absence of a readable book inside model weights would not automatically resolve whether copies made during data processing were legally significant.
The Lords committee’s 2026 report said that large-scale copying and processing during training may engage the reproduction right. It also noted that rightsholders often cannot determine whether their works were included because developers do not provide enough training-data transparency. The committee stated that the UK still had no definitive ruling on whether training a generative AI model on copyrighted works without a licence infringes copyright. Read the committee’s discussion of training and copyright.
Why creators and publishers object
Authors, publishers, journalists, photographers, musicians and other creators argue that training can transfer value from their work to model developers without permission or payment. Their concerns include:
- uncompensated copying may weaken markets that funded the original work;
- generated content may compete with licensed articles, stock images, music or other creative services;
- systems may reproduce recognizable passages or other protected expression;
- creators generally lack practical ways to verify whether their work was used;
- opt-out systems place the burden on individual rightsholders; and
- opaque datasets make attribution, compensation and enforcement difficult.
Those are serious objections, but lawsuits and allegations do not by themselves establish that every work used in every AI-training process was infringed. The claims involving organizations such as The New York Times, the Authors Guild and individual authors remain disputes over particular conduct, evidence and legal theories.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesWhat “fair use” does—and does not—answer
In the United States, fair use is a fact-specific doctrine, not a blanket exemption for AI training. Courts generally consider the purpose and character of the use, including commerciality and transformation; the nature of the copyrighted work; the amount and substantiality used; and the effect on actual or potential markets.
Rank #3
A commercial purpose can matter, but it does not automatically defeat fair use. Likewise, a model’s tendency to produce new text does not automatically make the underlying copying lawful. The analysis can differ depending on the dataset, the source works, the way copies were made, the model’s behavior and the effect on markets.
OpenAI’s later written evidence referred to two U.S. federal opinions that it characterized as finding AI training to be fair use. That statement should not be treated as a universal ruling covering every model, dataset or developer. It is OpenAI’s characterization of selected decisions, not a definitive answer to all AI-training disputes.
OpenAI’s strongest argument is economic, not just legal
OpenAI’s position has a practical market-access component. A licensing-only system could involve millions of rightsholders, fragmented ownership and substantial negotiation and reporting costs. Large technology companies may be better placed than startups to pay for broad catalogues of text, images, audio and code.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →OpenAI argues that mandatory licensing could therefore entrench companies that already have large budgets, content libraries or distribution platforms. The counterargument is equally important: a difficult business model does not create a legal entitlement to free inputs. If training commercially valuable systems depends on creative work, creators may reasonably argue that permission and compensation should be part of the model’s cost.
Rank #4
| OpenAI’s position | Rightsholders’ position | What is established |
|---|---|---|
| Broad access to contemporary data is needed for competitive models. | That access copies and monetizes creative work without permission. | The legality depends on jurisdiction, facts and legal theory. |
| Licensing could raise barriers for startups. | Free use can undermine licensing markets and creator incentives. | Both are genuine policy concerns; neither decides the legal question. |
| Training is different from distributing a verbatim work. | Training still requires copying and may enable substitution. | Courts must examine the full data and output pipeline. |
Where the UK debate stood in 2026
The original OpenAI submission came during an evolving UK policy debate. It should not be presented as the final UK position.
The Lords committee’s 2026 recommendations called for stronger licensing, transparency and enforcement rather than reforms that would remove incentives to license works for AI training. In May 2026, the committee said the government no longer had a preference for a broad copyright exception based on an opt-out mechanism and urged mandatory transparency requirements for large AI developers. See the committee’s recommendations and the May 2026 government-response announcement.
As of August 16, 2026, the most defensible conclusion is that the UK had not settled the core question. OpenAI had made a strong commercial and policy case for access to copyrighted training data, but its submission did not establish that all such copying is lawful, free or exempt from licensing obligations.
Why the Getty–Stability AI case did not settle training legality
The Getty–Stability AI litigation illustrates why individual court outcomes must be read narrowly. In the 2025 UK High Court proceedings, Getty abandoned its primary copyright claim after accepting there was no evidence that Stability AI’s model had been trained or developed in the UK.
Best Value
The court considered a separate issue involving whether model weights made available in the UK were themselves infringing copies. That claim failed, and permission to appeal was later granted on the secondary issue. The case therefore did not decide whether training a model on copyrighted works without a licence generally infringes the reproduction right.
A defendant can prevail because of territoriality, missing evidence, pleading problems or the specific legal theory being tested. Such an outcome is not the same as a declaration that all AI training is lawful.
What a workable compromise could involve
Possible approaches include:
- negotiated licences with publishers, stock libraries, music companies and rights-management organizations;
- collective or extended collective licensing;
- opt-in marketplaces for training data;
- mandatory training-data summaries, registers or audit records;
- machine-readable rights reservations;
- compensation funds or usage-based royalties;
- public-domain and openly licensed datasets;
- smaller domain-specific models trained on narrower licensed corpora;
- retrieval systems that access licensed databases at query time; and
- creator-controlled APIs with defined permissions and reporting.
None is a complete solution. A crawler-blocking system may limit future access but cannot necessarily compensate creators for past use. An opt-out mechanism is meaningful only if developers can detect and honor it. A certification service may provide compliance signals without granting a binding licence. Synthetic data can reduce dependence on original works but may reduce quality or amplify errors through recursive model training.
Any serious framework would need to answer what is licensed, where the permission applies, whether it covers training or retrieval, how use is audited, how compensation is calculated and which country’s law governs collection, training, deployment and distribution.
The bottom line on the headline
OpenAI did argue that leading modern models could not realistically be trained using only public-domain material. It also argued that training is legally distinct from reproducing a work for users and that broad licensing mandates could favor incumbents.
But “we cannot compete without access to this data” is not the same as “we have a right to copy it for free.” Public access does not erase copyright; a licence to access content is not necessarily permission to train a model; and a model’s lack of a human-readable copy does not end the copyright analysis. The decisive questions remain fact-specific, jurisdiction-specific and, in the UK as of August 2026, unresolved.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

