Tech

OpenAI and Microsoft warned of web ‘doom loop’ in internal files

Unsealed court documents reveal that the tech giants acknowledged their data scraping practices threatened the economic foundations of content suppliers, even as they pursued significant commercial gains.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: The Verge · View original source
OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
Markets & Finance

Recently unsealed court documents in the New York Times’ lawsuit against OpenAI and Microsoft have revealed that both companies internally warned their data scraping practices were creating a “doom loop” for the web. The internal files describe the harvesting of data as the “largest theft of labor in human history” and a “complete mockery” of the concept of fair use, suggesting a level of awareness regarding the impact on publishers that contradicts earlier public statements.

Microsoft has attempted to distance itself from the comments of its Director of Applied Science, Brent Hecht, who characterised the data scraping in such stark terms. Spokesperson Alex Haurek stated that the comments reflect an individual perspective rather than company views. However, in a separate court filing, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, characterised Hecht’s role as adversarial, noting that he holds “divergent, academic, and forward-looking views” and is employed to bring asymmetrical points of view to the table.

Despite these attempts to manage the narrative, internal Microsoft documentation explicitly stated that the AI content strategy had started a “doom loop” that would hurt model performance and the entire web simultaneously. The document noted that it is highly unusual for an end-product to threaten the economic foundations of its essential suppliers, yet that is the situation created for the large language model business with respect to its content supply chain.

The documents also highlight the shift in user behaviour, with Microsoft CEO Satya Nadella admitting that chatbots have largely replaced search, removing the need for users to go directly to source material. OpenAI’s Head of ChatGPT, Nick Turley, stated that once a user receives an answer from the chatbot, there is “no good reason to click” on a link to the source. This dynamic has been described by Microsoft as a product that destroys its own supply chain.

Financial implications are significant, with OpenAI media and economic experts attributing a drop in referral traffic for sites like the New York Times directly to AI summaries. They speculated that search referrals may be down as much as 60 per cent. Meanwhile, OpenAI cofounder Greg Brockman expressed interest in the “gazillions” of dollars potentially made through commercial AI, even as the company acknowledged that GPT-4 “memorised a ton of data” and would be “insanely good at regurgitation.”

The New York Times argues that these internal warnings demonstrate that both companies knew their AI products were threatening the economic foundations of content suppliers yet proceeded in pursuit of significant commercial gains. The unsealed documents provide a detailed account of the internal debates and admissions that have now become central to the copyright infringement case.

Continue reading

More from Tech

Read next: Anthropic appoints Accenture as first embedded AI safety evaluator
Read next: Sony and UMG sue Suno over alleged copyright infringement in new AI models
Read next: Clicks Communicator shipments to start in December as price rises to $649