Publishers and authors sue Google over alleged unauthorised AI training data
Hachette, Cengage, Elsevier, and author Scott Turow accuse Google of training its Gemini platform on copyrighted material without permission, citing internal documents that warn of fines up to $100 billion.

A coalition of major publishers and authors has initiated a class action lawsuit against Google in the U.S. District Court for the Southern District of New York, alleging the technology giant trained its Gemini artificial intelligence platform on copyrighted works without authorisation. The plaintiffs, comprising Hachette, Cengage, Elsevier, author Scott Turow, and the S.C.R.I.B.E. organisation, argue that Google systematically removed or altered copyright information to conceal the use of these materials.
The complaint centres on the allegation that Google utilised copies of works sourced from its own Google Books service and the Google Play store for AI training purposes. The lawsuit contends that these entities provided Google with access to copyrighted content specifically for the purpose of enabling book searches, a service that restricts users to viewing only short snippets and bibliographic data rather than full texts. The plaintiffs assert that Google exceeded this limited scope by incorporating the full works into its AI models without seeking the necessary permissions.
In support of its claims, the filing references an internal Google document which allegedly warns that using copyrighted books for AI training could be “highly problematic” and expose the company to potential fines ranging from $10 billion to $100 billion. The lawsuit states that Google knowingly copied these works from scope-limited programs, fully aware it lacked the authorisation to do so. Google did not immediately respond to a request for comment regarding the allegations.
This legal action arrives against a backdrop of intensifying litigation between the creative industry and AI developers. While two recent court decisions in California have favoured AI companies by ruling that such training constitutes “fair use” under existing U.S. copyright law, the plaintiffs in the New York case hope for a different judicial perspective. The conflict remains nuanced, with previous rulings not necessarily establishing an inarguable precedent for all jurisdictions.
The lawsuit also highlights the broader context of copyright disputes in the AI sector, including a separate $1.5 billion settlement involving Anthropic, which marked the largest payout in U.S. copyright law history. Although approximately half a million writers were eligible for payments of at least $3,000 under that settlement, many opted out to pursue further legal action. The current filing against Google underscores the ongoing tension between tech companies’ reliance on vast datasets and the rights of publishers and authors to control the commercial use of their intellectual property.
