Tech

Publishers and authors sue Google over alleged unauthorised AI training data

Hachette, Cengage, Elsevier, and author Scott Turow accuse Google of training its Gemini platform on copyrighted material without permission, citing internal documents that warn of fines up to $100 billion.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · original
Google faces another AI training lawsuit from major publishers
Class action filed in New York alleges tech giant stripped copyright metadata to hide use of stolen works

A coalition of major publishers and authors has initiated a class action lawsuit against Google in the U.S. District Court for the Southern District of New York, alleging the technology giant trained its Gemini artificial intelligence platform on copyrighted works without authorisation. The plaintiffs, comprising Hachette, Cengage, Elsevier, author Scott Turow, and the S.C.R.I.B.E. organisation, argue that Google systematically removed or altered copyright information to conceal the use of these materials.

The complaint centres on the allegation that Google utilised copies of works sourced from its own Google Books service and the Google Play store for AI training purposes. The lawsuit contends that these entities provided Google with access to copyrighted content specifically for the purpose of enabling book searches, a service that restricts users to viewing only short snippets and bibliographic data rather than full texts. The plaintiffs assert that Google exceeded this limited scope by incorporating the full works into its AI models without seeking the necessary permissions.

In support of its claims, the filing references an internal Google document which allegedly warns that using copyrighted books for AI training could be “highly problematic” and expose the company to potential fines ranging from $10 billion to $100 billion. The lawsuit states that Google knowingly copied these works from scope-limited programs, fully aware it lacked the authorisation to do so. Google did not immediately respond to a request for comment regarding the allegations.

This legal action arrives against a backdrop of intensifying litigation between the creative industry and AI developers. While two recent court decisions in California have favoured AI companies by ruling that such training constitutes “fair use” under existing U.S. copyright law, the plaintiffs in the New York case hope for a different judicial perspective. The conflict remains nuanced, with previous rulings not necessarily establishing an inarguable precedent for all jurisdictions.

The lawsuit also highlights the broader context of copyright disputes in the AI sector, including a separate $1.5 billion settlement involving Anthropic, which marked the largest payout in U.S. copyright law history. Although approximately half a million writers were eligible for payments of at least $3,000 under that settlement, many opted out to pursue further legal action. The current filing against Google underscores the ongoing tension between tech companies’ reliance on vast datasets and the rights of publishers and authors to control the commercial use of their intellectual property.

Continue reading

More from Tech

Read next: Open-source tool claims 97 per cent token savings for AI agents
Read next: Valvoline Unveils August 2026 Promotional Offers for Service and Retail Buyers
Read next: Developer Antirez releases native MiniMax H3 inference engine for Apple Silicon