Three publishers, novelist Scott Turow and his company SCRIBE have filed a proposed class-action lawsuit accusing Google of copying millions of books and journal articles to form Gemini. This includes works provided through Google Books, Play Books and Scholar.
On July 10, Hachette Book Group, Cengage Learning, Elsevier, Turow and SCRIBE came together to a lawsuit filed in the United States District Court for the Southern District of New York, with the Association of American Publishers announce it the same day. They argue that the books and articles provided to these services were intended for specific purposes, and that their use to train a commercial AI model was not one of them. The lawsuit also claims that Google copied works obtained from pirate websites and paid libraries. Google did not comment on the complaint at the time of publication, and no court has ruled on any of the claims. The main question is whether the permission for use also covers training a model on this data.
What the complaint claims
The complaint contains four counts. Three of them allege unauthorized reproduction under copyright law, covering Google Books and other Google services, web scraping uploads and copying during training. The fourth alleges that Google removed copyright management information in violation of the DMCA. The plaintiffs seek damages, an injunction, a detailed accounting of the works used by Gemini for training, and court orders to remove any unauthorized copies. The filing cites what it describes as internal Google documents, one of which calls the use of Google Play Books for AI “highly problematic for Google,” with potential fines ranging from “$10 billion to $100 billion.” He attributes another line to Gemini’s chief engineer, who reportedly told colleagues: “We don’t make deals for data we already have or already own.” None of these documents are public and the citations come from the plaintiffs’ files.
Where the robot controls stop
Google-Extended is the robots.txt token that covers the content that Google crawls from your site. This limits the ability to use this content for future Gemini training and some basic uses. Neither provisioning method discussed here involves this token. The books were provided directly to Google via agreements, so a robots.txt file does not affect this process. The web-scraping allegations refer to copies that the complaint says appeared in Common Crawl after being hosted on pirate sites and subscription libraries. Since these copies are hosted on different domains, a robots.txt file cannot regulate them.
On June 25, Google published a political document arguing that training on public web data is a “transformative, non-expressive use” under fair use protections. The document also mentions machine-readable controls, like Google-Extended, that websites can use to opt out. However, the documents examined here would have arrived through different channels.
Last month, Digital content Next sent a cease and desist letter to the Common Crawl Foundation, arguing that copyright law does not function as an opt-out system.
Why it matters
The issue of permission and fair dealing are separate concerns. Fair use may apply even if there is no agreement authorizing the use, and the complaint resolves neither issue.
Your crawler settings are a factor smaller than this situation might imply. In January, BuzzStream data reported that 79% of major news sites block at least one AI training bot, which is the Google-Extended address channel. The two groups of copies analyzed here would have passed through routes that these parameters do not affect.
Looking to the future
In 2025, two Northern California decisions found that the uses of training at issue were fair based on the records available to them. The anthropogenic courtyard denied summary judgment on the pirated copies of the central library, while Judge Meta stressed its decision was specific to these plaintiffs and their case. The editors said They filed the suit in New York after initially planning to intervene in Google Generative AI’s ongoing copyright litigation in California, and that the new suit preserves claims that they say do not fall within this proposed class. The next step is Google’s response, either a response or a motion to dismiss.





