OTWopentechwire
Tech Intelligence, Openly Wired
Policy

Microsoft Scientist Labelled OpenAI Training Methods "Largest Theft of Labour in Human History"

Unsealed court filings reveal executives at both companies recognised AI scraping as an existential threat to journalism, even as legal teams defend the practice

KW
Kenji Watanabe
Hardware & Products Reporter · Tokyo
Sep 21, 2026
6 min read
Microsoft Scientist Labelled OpenAI Training Methods "Largest Theft of Labour in Human History"
Microsoft Scientist Labelled OpenAI Training Methods "Largest Theft of Labour in Human History"Credit: Shutterstock

A "Doom Loop" in Court Documents

Dr Brent Hecht, a director of applied science at Microsoft, described OpenAI's web scraping operation as the "largest theft of labour in human history" and warned it could trigger a "doom loop" for the publishing industry. The comment appears in newly unsealed court filings from a 2023 lawsuit filed by a major US newspaper against both companies.

At Opentechwire, we've tracked the legal battles over training data since the first wave of generative AI products reached consumers. This disclosure stands out because the warnings came from inside the room. Nick Turley, an OpenAI executive, used the term "existential threat to publishers" in internal communications. Another Microsoft document from 2023 stated that millions worldwide would soon view large models as "hoovering up" all their work in "an astonishing theft of unprecedented proportions."

The documents were released as part of discovery in ongoing copyright litigation. Only excerpts have been made public, without surrounding context. Yet the fragments paint a picture of companies whose technical and business staff understood the stakes, even as their legal strategies moved in a different direction.

Paywalls, Copyright Notices, and "Ah Nice"

The filings detail specific methods used to acquire training material. OpenAI and its partners reportedly bypassed paywalls on news sites, scraped millions of documents, and removed copyright notices from the data fed into large language models.

One exchange shows an employee informing OpenAI president Greg Brockman about a new paywall circumvention technique. Brockman's reply, according to the documents: "ah nice." The brevity of the response contrasts with the scale of the operation it greenlit.

These revelations arrive at a moment when several similar lawsuits have been dismissed on procedural grounds, with judges noting that copyright law has not yet adapted to generative AI. The case involving the newspaper consortium is widely seen as a bellwether. If courts rule that scraping constitutes fair use, the training-data question may be settled for years. If they do not, the industry faces a reckoning over licensing and compensation.

Supply Chain Destruction

Dr Hecht's framing is particularly sharp. He described large AI models as "a product that destroys its supply chain." The metaphor is precise: if users can get summaries, rewrites, or synthesised answers from a chatbot, traffic to the original publishers evaporates. An OpenAI software engineer acknowledged this in 2023, telling colleagues that "no matter how prominently we show the links, users won't click." Turley confirmed the dynamic, noting that AI products are "largely substitutive" to journalism.

This is not a hypothetical concern. Publishers across Asia, Europe, and North America have reported measurable declines in referral traffic since late 2022, when ChatGPT launched. Some have struck licensing deals with OpenAI and other model builders; others have blocked crawler bots. A few have done both, betting that a hybrid strategy offers the best hedge against an uncertain future.

The internal Microsoft and OpenAI comments suggest that executives foresaw this outcome. Whether they believed the trade-off was justified, or simply inevitable, remains unclear from the released excerpts.

Microsoft Distances Itself from Its Own Scientists

Microsoft moved quickly to distance itself from the statements made by its employees. A spokesperson told reporters that the company's position is set out in its court filings, which argue that these uses are transformative and consistent with copyright law. The statement emphasised that Microsoft's Copilot product is not a substitute for publishers' journalism.

That assertion sits awkwardly alongside the internal documents. If senior applied scientists and strategy staff believed the opposite, it raises questions about alignment between technical assessment and legal argument. It also complicates Microsoft's broader pitch to enterprise customers, many of whom operate in sectors where intellectual property and data provenance matter.

The dynamic is not unique to Microsoft. Across the generative AI sector, companies have adopted a public stance that training on web-scraped data is legal and beneficial, while internal discussions acknowledge risks that range from copyright liability to the collapse of content ecosystems. The gap between these two narratives has widened as litigation has intensified.

The Fair Use Question

At the heart of the lawsuit is whether large-scale scraping and training constitute fair use under US copyright doctrine. Fair use permits limited reproduction of copyrighted material for purposes such as criticism, commentary, news reporting, teaching, and research. Courts weigh four factors: the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market for the original.

AI companies argue that training is transformative, that it does not republish the original works, and that it serves a public benefit by advancing technology. Publishers counter that the scale of copying is total, that the models can generate text that substitutes for their articles, and that the market harm is already visible in declining traffic and advertising revenue.

So far, judges have been reluctant to rule definitively. Several cases have been dismissed on standing or pleading issues, with courts noting that the law is unsettled. That makes the current lawsuit especially important. A decision on the merits would set a precedent, and the newly disclosed internal communications could influence how judges assess the intent and impact of the scraping operations.

What Happens If the Supply Chain Breaks

Dr Hecht's "doom loop" concept warrants unpacking. If AI-generated summaries displace traffic to news sites, advertising revenue falls. Newsrooms shrink, original reporting declines, and the volume of high-quality material available to train future models diminishes. Over time, the models degrade or stagnate, and the AI products become less useful. The loop closes.

This scenario is not universally accepted. Some researchers argue that models can be fine-tuned on synthetic data or domain-specific corpora, reducing reliance on fresh journalism. Others point out that many news organisations have survived previous disruptions, from free classifieds sites to social media aggregation. The question is whether this disruption is different in kind or simply in speed.

At Opentechwire, we've observed that the debate often splits along geographic lines. In markets with strong public broadcasters or state-supported media, the existential threat feels less acute. In advertising-dependent ecosystems, especially in Southeast Asia and parts of Latin America, the concern is immediate. Publishers there have less room to negotiate licensing deals and fewer resources to pursue litigation.

The Timing of the Disclosures

The documents were unsealed in September 2026, more than three years after the lawsuit was filed and nearly four years after ChatGPT's public launch. The delay reflects the pace of discovery and the sensitivity of the material. Both Microsoft and OpenAI sought to keep the communications under seal, arguing they contained proprietary information and could prejudice the case.

The timing also matters because the AI landscape has shifted. In 2023, when many of these internal messages were written, the industry was still debating whether generative models would find mainstream adoption. By 2026, that question is settled. The focus has moved to regulation, safety, and economics. The unsealed documents pull the conversation back to a foundational issue: who owns the data, and who should benefit from its use.

What Comes Next

The lawsuit is expected to proceed to trial in 2027, barring a settlement. Both sides have signalled they are prepared for a lengthy fight. The plaintiffs have added claims and amended their filings; the defendants have filed motions to dismiss and compel arbitration in some instances.

Meanwhile, parallel cases are moving through courts in the United Kingdom, the European Union, and Canada. Each jurisdiction applies different copyright frameworks, and the outcomes may diverge. A ruling in favour of publishers in one region could create pressure for licensing deals elsewhere, even if other courts rule differently.

For now, the disclosed comments serve as a reminder that the people building these systems understood the risks. Whether that knowledge translates into legal liability, or simply into awkward testimony, remains to be seen.

Read next
Policy

Singapore and Malaysia Pull Ahead as AI Investment Widens ASEAN's Economic Divide

Mei-Lin Tan · 5 min
Policy

Why Tech CEOs Keep Asking to Regulate Themselves

Priya Nair · 5 min
Policy

Beijing Signals Retaliation Over Washington's AI Distillation Accusations

Mei-Lin Tan · 5 min
Spot something wrong? Email corrections@opentechwire.com. We log every correction publicly.