Reported by 1 source

The short version

  • Internal memos describe the harvesting of web content as creating a doom loop that threatens the economic viability of publishers.
  • Microsoft has distanced itself from an employee's characterization of the data usage as the largest theft in human history.
  • Executives acknowledged that AI models could replace search engines and substitute for the labor of cultural creators.

Recently unsealed court documents in the ongoing legal dispute between The New York Times and artificial intelligence developers OpenAI and Microsoft reveal stark internal warnings about the industry's trajectory. The filings contain documentation suggesting that executives at both companies were aware their strategies posed significant risks to the digital ecosystem. Internal communications described the rapid expansion of large language models as initiating a destructive cycle that would undermine the very sources feeding these systems.

A central theme in the released materials is what one internal Microsoft document termed a doom loop. This concept describes a scenario where AI products threaten the economic foundations of their essential suppliers. The text notes that it is highly unusual for an end product to endanger its content supply chain, yet this is precisely the situation created by the large language model business. The warning implies that as AI models become more prevalent, they may degrade the quality and availability of the web content required to train future iterations.

News Journal

Brent Hecht, Microsoft’s Director of Applied Science, authored several of the most critical assessments found in the filings. He characterized the wholesale scraping of data used to train models as an unprecedented appropriation of effort. His internal notes suggested that the defense of such practices under fair use doctrines was fundamentally flawed. Hecht argued that the scale of data harvesting represented a historic shift in how intellectual labor is valued and utilized by technology firms.

Microsoft has moved quickly to distance itself from Hecht’s specific language and conclusions. A company spokesperson stated that these comments reflect an individual employee’s perspective rather than official corporate views or legal analysis. In separate court filings, Jordan Usdan, a general manager at Microsoft AI, described Hecht’s role as providing adversarial and forward-looking academic viewpoints. Usdan emphasized that Hecht is employed to offer asymmetrical perspectives on data ecosystems and does not speak for the company regarding theoretical impacts on content creators.

Despite efforts to contextualize these internal warnings, other documents suggest a broader awareness of the risks among leadership. Satya Nadella, Microsoft’s CEO, reportedly acknowledged that chatbots have largely replaced traditional search functions. This shift removes the necessity for users to visit original sources for information, potentially diverting traffic and revenue away from publishers. The admission highlights a fundamental change in user behavior driven by AI integration into daily digital interactions.

OpenAI’s internal communications also reveal concerns about copyright and data integrity. Employees noted that earlier models had memorized significant amounts of training data, making them highly effective at reproducing copyrighted material verbatim. While the company stated that preventing such memorization was important to minimize legal violations, internal assessments admitted that the technology remained prone to regurgitating text from sources like The New York Times and other major publications. This capability raises questions about the distinction between synthesis and reproduction.

The financial motivations behind these developments are also scrutinized in the filings. Contrary to public narratives emphasizing altruistic goals, documents indicate a strong focus on commercial potential. Greg Brockman, an OpenAI co-founder, was noted for his interest in the substantial profits that could be generated through commercial AI applications. This profit motive appears to have driven aggressive data acquisition strategies, even as internal voices raised ethical and legal concerns about the methods employed.

The tension between stated principles and actual practices is evident in discussions around paywalled content. While Nadella has publicly suggested that paywalled material should be licensed for use, an OpenAI representative admitted to being unaware of any systematic efforts to detect or remove such content from training datasets. This discrepancy suggests a gap between high-level policy statements and the operational realities of data collection at scale.

Jack Clark, OpenAI’s Policy Director, also contributed to the internal discourse by noting that AI systems were effectively substituting for human labor in cultural production. He observed that these technologies were replacing the work of individuals who define societal culture. This perspective aligns with broader concerns about the displacement of creative professionals and the devaluation of original content in an AI-driven economy.

As the legal proceedings continue, these documents provide a window into the strategic calculations made by leading tech firms. They illustrate a conflict between the drive for innovation and the preservation of existing economic structures. The outcome of this case may set important precedents for how data is sourced, used, and compensated in the age of artificial intelligence.

Sources behind this briefing

Go to the original reporting

  • The Verge↗OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web