Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A study co-authored by O’Reilly Media CEO Tim O’Reilly found that OpenAI’s GPT-4o distinguished tested passages from paywalled O’Reilly books better than chance. The result is consistent with prior exposure to those passages, but it does not prove that OpenAI copied the books directly, obtained them unlawfully, or infringed copyright.

What O’Reilly and the researchers alleged

On April 1, 2025, the AI Disclosures Project released a working paper by Tim O’Reilly, Ilan Strauss and Sruly Rosenblat. It reported evidence consistent with GPT-4o having encountered non-public material from O’Reilly Media books. The paper tested model behavior; it did not uncover an OpenAI training-data list, internal document or record of a paywall being bypassed. The Social Science Research Council’s announcement describes the paper and its release.

O’Reilly has also argued more broadly for disclosure and compensation when AI companies use copyrighted work commercially. His company’s stated approach calls for licensing, attribution and ways to track and compensate creators. O’Reilly’s position on generative AI is relevant context: he is both a co-author of the study and the publisher whose books were tested.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the study tested

The researchers examined 34 copyrighted O’Reilly books and 13,962 paragraph excerpts, according to reporting on the study. They compared publicly accessible portions with passages treated as paywalled and tested OpenAI’s GPT-4o, GPT-3.5 Turbo and GPT-4o Mini. The paper used DE-COP, a membership-inference technique designed to detect whether a model responds differently to original human-written text than to paraphrased or AI-generated alternatives. The paper’s arXiv record identifies the book sample; the working paper describes the method.

In broad terms, the test looks for a stronger model signal for an original passage than for comparison text. Better-than-chance discrimination can suggest that a model previously encountered the passage. This is a statistical inference, not a simple request for the model to recite a book, and it is not a direct inspection of the model’s training corpus.

What the reported numbers mean

The AI Disclosures Project reported an AUROC of about 0.82 for GPT-4o on paywalled passages. AUROC measures how well a classifier separates two categories across decision thresholds: 0.50 is roughly chance performance, while 0.82 indicates stronger-than-chance discrimination in this test. It does not mean that 82% of the books were copied, that 82% of GPT-4o’s training data came from O’Reilly, or that 82% of the passages were memorized. The project reported a bootstrapped 95% confidence interval of 0.60–0.96, a wide range that cautions against treating 0.82 as a precise estimate. The project’s results summary gives the score and interval.

GPT-4o Mini was near chance in the reported tests: the project summary gives roughly 0.56, while the paper’s broader characterization is approximately 0.50. GPT-3.5 Turbo showed relatively greater recognition of public excerpts than paywalled ones. These differences matter: the reported signal was not uniform across the tested models, and the GPT-4o result should not be generalized to every OpenAI model, particularly models released or offered after the 2025 study.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why compare public and paywalled passages?

O’Reilly books are not uniformly inaccessible. Some material, such as selected chapters or excerpts, has been publicly available, while other content is offered through its subscription service. O’Reilly describes a framework that treats selected portions as public and others as paywalled in its discussion of copyright-aware AI. That account helps explain the comparison: public passages could plausibly have entered a model from the open web, while stronger recognition of paywalled passages may point to another route of exposure.

But “paywalled on O’Reilly’s site” does not mean a passage could not have appeared elsewhere. It might have been available through a preview, library or licensed database, search index, user submission, pirated copy, or another intermediary. The comparison makes the result more informative than a test with no access distinction, but it cannot establish where the model encountered the text.

What the evidence does—and does not—establish

  • It supports: GPT-4o showed a behavioral pattern consistent with prior exposure to the tested paywalled passages.
  • It does not identify: the source, date or route by which any passage reached the model, or the person or organization that supplied it.
  • It does not establish: that OpenAI directly downloaded O’Reilly’s subscription library, bypassed its paywall, copied the books verbatim into training data, or committed copyright infringement.
  • It does not cover: every O’Reilly title, every OpenAI model, or books from other publishers.

The paper’s authors acknowledge alternative routes, including text supplied by users who pasted passages into ChatGPT. Other possibilities include licensed databases, public or semi-public copies, aggregators, publisher-provided material, or pirated sources. Recognition can also reflect memory of exact wording or distinctive phrases, or patterns induced by the test; it is not by itself proof that a complete, retrievable book is stored in the model. TechCrunch’s report discusses the source uncertainty and user-submitted text possibility.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What OpenAI has said publicly

OpenAI’s GPT-4o system card says the company forms partnerships to access some non-public data, including paywalled content. That general statement does not identify O’Reilly books, demonstrate an agreement with O’Reilly, or explain the origin of the passages in this experiment. The system card is not a specific response to the paper. The materials cited here do not show a detailed, issue-specific OpenAI rebuttal or confirmation.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Coverage of the paper says O’Reilly did not have a licensing agreement with OpenAI. Even if there was no direct publisher agreement, that alone would not establish that OpenAI had no other lawful source or permission. The study does not resolve the source or authorization question.

Why the copyright question remains open

Several issues are often compressed into the word “stole,” but they are distinct. The tested works are copyrighted; some passages were treated as paywalled on O’Reilly’s service; the model showed a recognition signal; and the paper does not establish the route by which the passages became available. Whether a particular use was authorized, whether copying occurred in a legally significant way, and whether that use infringed copyright require evidence and legal analysis beyond this experiment, including the relevant jurisdiction and facts about acquisition and use.

That distinction does not make the finding trivial. A reproducible method for identifying possible exposure could help scrutinize opaque training practices, especially if it is independently replicated across publishers and models. But the result is an audit signal, not a chain-of-custody record or a court finding.

What could settle more of the dispute

The claim would be stronger with independent replication under a preregistered protocol, tests by researchers without a direct institutional or commercial relationship to O’Reilly, and evidence that the tested passages were not accessible through other channels. Training-data documentation, vendor records or discovery evidence could help establish provenance. Conversely, evidence of widespread public availability or user submission, a documented license, or failed independent replications could weaken the inference that the signal reflects exposure through an unauthorized training source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For publishers and authors, the practical issue is whether AI developers will disclose the provenance and permissions behind training data and offer workable licensing and compensation terms. O’Reilly advocates such mechanisms, but the study does not establish a particular licensing model or decide what the law requires.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.