"Sketchy AF"

How OpenAI got its books for its AI training

Newly unsealed court documents from the consolidated New York litigation against OpenAI and Microsoft show how early those involved knew where their book data came from. Staff called the source dubious in internal messages, discussed how to describe it publicly, deprioritised buying books as too expensive, and later deleted the datasets partly for legal reasons.

When the DIGITAL PUBLISHING REPORT covered Project Panama in August, the question was why AI companies are now buying second-hand non-fiction in bulk, scanning it and destroying it. The documents that have now become public provide the backstory. They show how one of the leading AI companies obtained books before the detour through the antiquarian trade.

Books1, Books2 and a shadow library

The documents come from the consolidated case In re OpenAI before the U.S. District Court for the Southern District of New York. The plaintiffs include the Authors Guild and writers such as John Grisham, David Baldacci, Jodi Picoult and Jonathan Franzen, alongside several news publishers including the New York Times. Both sides have filed motions for summary judgment. The authors, who filed on 4 September, are asking the court to rule that OpenAI and Microsoft infringed their copyrights and to reject the fair-use defence. The court has not yet ruled. The motions were initially filed in redacted form and have since been partially unsealed; they will determine which claims over training practices proceed to trial.

A court order issued in February had already established that an OpenAI employee downloaded pirated copies of books from Library Genesis (LibGen) in 2018. These became the "Books1" and "Books2" datasets, with around 12 and 55 billion tokens respectively, which were used to train GPT-3 and GPT-3.5. According to OpenAI's lawyers, their use for training ended in late 2021, and the data was deleted in mid-2022.

What is new is how openly the source was discussed internally. According to the plaintiffs' filing, as also reported by the Wall Street Journal, OpenAI engineers Tom Brown and Ben Mann described LibGen as "sketchy AF". Dario Amodei, then OpenAI's research director, asked in an internal message whether it was "sketchy" to call the corpora Books1 and Books2 without saying what they were, given that they actually came from LibGen. Mann described a public-facing account of the training data as "deliberately vague since it's LibGen". Meeting notes from April 2020 cited in the filing called for an upcoming paper with "no code, vague on data". The GPT-3 paper published in May 2020 did in fact refer to Books1 and Books2 without naming LibGen. Staff also noted internally that the use conflicted with copyright restrictions, yet, according to the plaintiffs, tried to conceal the piracy. Researcher Sam McCandlish said his concern was less about the law than about optics, specifically a Hacker News headline about OpenAI using copyrighted data from a dubious Russian website.

The cast of characters is notable. Amodei, Mann, Brown, McCandlish and Jack Clark, who is also quoted in the filings, are now among the founders and leadership of Anthropic. That company agreed to a $1.5 billion settlement in Bartz v. Anthropic over its own downloads from shadow libraries.

Deleted "for legal reasons"

Why the datasets disappeared is also contested. OpenAI had told the court that Books1 and Books2 were deleted because they were no longer in use. The plaintiffs, however, cite a June 2022 message from then vice-president of research Bob McGrew saying that, given OpenAI's growing public profile, it was the right time to remove LibGen from its systems, and that doing so would be "very valuable for legal reasons". Internally, the deletion effort was known as "Project Clear".

Expensive per token

One detail in the filings is particularly telling for the publishing industry. In late 2022, OpenAI president Greg Brockman wrote to co-founder Ilya Sutskever that the company could buy books in large quantities, but that they were "expensive per token" and had therefore not been prioritised. What OpenAI deprioritised as uneconomical back then has become industry practice since the Bartz ruling. Lawfully buying physical books is now seen as the legally safer route, while pirated digital copies have turned into a costly liability.

The filings suggest little restraint with news content either. When staff told Brockman about a workaround for the New York Times paywall to support scraping, he responded approvingly.

Microsoft knew early on

According to the plaintiffs, Microsoft knew about the use of LibGen as early as April 2019, when Sam Altman and Dario Amodei presented an early version of GPT-3 to Bill Gates and disclosed the use of LibGen to Gates and Microsoft CTO Kevin Scott, among others. Microsoft CEO Satya Nadella testified that downloading pirated content is clearly illegal and that paywalled content should be licensed for AI training. Had he known that OpenAI scraped and trained on paywalled content, he said, he would have required OpenAI to retrain its models. A Microsoft spokesperson said Nadella's testimony and the company's position in the case were consistent.

The filings also describe exchanges between the two companies. Through "Project Taxi", OpenAI received billions of web pages that Microsoft had collected for its Bing search engine. Under "Project Mango", OpenAI paid Microsoft to build and operate a crawler. OpenAI also passed copies of books to outside contractors, including for a project in which novels were read and summarised, and set up a central library that included books not used for training. That point is likely to interest the plaintiffs, since in Bartz it was precisely the central library built from pirated copies that the court refused to treat as fair use.

Substitute, not supplement

Authors Guild CEO Mary Rasenberger observed, “These filings reveal shocking disdain for writers and their work through repeated, intentional decisions to steal books rather than pay for them with full knowledge that their products will destroy the careers of authors. OpenAI pursued its mass piracy scheme even in the face of clear evidence that doing so would degrade American culture by substituting human works with AI slop and would put thousands of writers out of work. We are grateful to our entire legal team for uncovering and telling this critical story.”

Statements about market impact could carry the most legal weight, as they bear on the fourth fair-use factor, the effect on the market for the original work. In 2020, Jack Clark, then policy director at OpenAI, wrote: "The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon." He added that OpenAI's work in this area would "make people unemployed". Tarun Gogineni, who joined OpenAI in 2022 to work on the models' writing quality, reportedly characterised potential displacement as "acceptable economic disruption". According to the news publishers, OpenAI executive Nick Turley wrote: "Our products are largely substitutive, period."

Microsoft's director of applied science Brent Hecht described the unauthorised copying of millions of news articles internally as possibly the biggest theft of labour in human history. Microsoft stresses that this reflects one employee's personal view and is not a legal analysis. The news publishers back their argument with Microsoft data cited in their motion, according to which click-through rates for Times and Daily News content fell by 83 to 93 per cent, and by 51 to 94 per cent for Ziff Davis domains.

OpenAI's position

OpenAI told the Wall Street Journal that the employees involved in creating the LibGen-derived datasets are no longer with the company and that the data was not used to develop its current ChatGPT models. The company disputes the authors' claims and maintains that training is a transformative, non-expressive analytical use that neither substitutes for nor harms the market for the original works. It has support from Washington, as the Trump administration has urged the court to rule that training AI on copyrighted works is fair use.

Analysis

The pattern echoes Bartz v. Anthropic. There, Judge William Alsup separated the training itself, which he held to be fair use, from building a library of pirated copies, which was not covered. The $1.5 billion settlement that followed received final approval from Judge Araceli Martínez-Olguín in July 2026; appeals against the approval have since been filed. The New York court is not bound by Alsup's reasoning. If it follows it, the LibGen question moves to the centre of the case against OpenAI, and internal messages about deliberate vagueness and deletions for legal reasons could matter when it comes to wilfulness.

For European publishers, too, the case is far from abstract. In Germany, Carlsen Verlag, author Marc-Uwe Kling and illustrator Astrid Henn are suing OpenAI Ireland before the Munich Regional Court I. They allege that ChatGPT generates stories and illustrations that closely copy the children's book series "Das NEINhorn", down to complete print templates with the Carlsen logo and an invented ISBN. The New York filings set no legal standard for that case, but they offer a picture of how deliberately decisions about training data were made.