Newsclip — Social News Discovery

Business

AI's Data Harvesting: A Crisis of Labor and Ethics

September 17, 2026
  • #Artificialintelligence
  • #Copyrightlaw
  • #Ethicsintech
  • #Journalism
  • #Airegulation
0 views0 comments
AI's Data Harvesting: A Crisis of Labor and Ethics

The Unvarnished Truth Behind AI Training

When I first started following the AI industry's rapid evolution, I thought we were witnessing a technological revolution that would benefit everyone. What I've learned in recent months is that the true story behind artificial intelligence's rise is more complex—and troubling—than we imagined. New unredacted court filings in The New York Times vs. OpenAI and Microsoft expose an uncomfortable reality: the companies we've come to trust for innovation are quietly engaging in practices that some of their own executives consider tantamount to theft.

Microsoft's Internal Admission

One of the most telling admissions came from a top Microsoft executive, who privately described the AI training process as "theft." This wasn't just an emotional outburst. It was a formal acknowledgment that these firms were essentially appropriating intellectual property without consent, often through means that circumvented existing paywalls and protections. The implications are staggering—this isn't just about data; it's about labor.

What Was Scanned?

The scale of content being scraped is unprecedented. OpenAI's mid-training datasets alone contain over 91,000 copies of works from publishers like The New York Times and the Daily News. A single Common Crawl-derived dataset included more than 2 million documents specifically from nytimes.com. This is not just a few articles; it's a wholesale extraction of content that publishers have spent years producing and monetizing.

"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers," read one internal Microsoft document, referring to how their AI models were harming the very publications that fed them.

The Human Cost

This isn't just a legal or ethical quandary—it's a human issue. As Microsoft's own data shows, the use of generative AI like Copilot has led to significant drops in click-through rates for The Times' content, with one report noting a 93% decline. These numbers don't reflect merely lost revenue—they point to real consequences for journalism and its workforce.

Microsoft's Director of Applied Science, Brent Hecht, went so far as to call this "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." This is a stark warning about what happens when we fail to properly value the contributions of creators and journalists who produce the very data that powers our AI future.

Legal Loopholes and Ethical Dilemmas

The legal argument surrounding fair use has been a contentious issue, with courts often favoring AI firms' arguments. However, these newly revealed admissions seem to contradict the premise that AI training is non-infringing, particularly when it comes to substitution and market harm.

  • OpenAI's own leadership recognized an "existential threat" to publishers
  • Microsoft CEO Satya Nadella testified that anything behind a paywall should be licensed for use
  • Internal documents reveal that the companies went to great lengths to circumvent paywalls
  • Copies of articles were stripped of copyright notices before training data was fed into models

This isn't just about money—it's about control over information and the rights of content creators. The fact that these internal communications were so candid speaks volumes about how little oversight exists in this rapidly growing sector.

What Happens Next?

As I continue to analyze developments like these, I'm reminded that we're standing at a crossroads. On one side, we have powerful AI firms pushing boundaries and redefining what's possible. On the other, we have publishers, journalists, and creators who deserve compensation and recognition for their work.

These filings are not just documents in a lawsuit—they're a wake-up call. They show us how far we've come—and how far we still need to go in terms of ethical data use and responsible innovation. We must demand transparency, accountability, and a reevaluation of the current system that allows for such large-scale harvesting of labor.

Final Thoughts

The rise of AI has brought incredible benefits, but we must also confront the real human costs behind it. These revelations from Microsoft and OpenAI force us to reconsider how we value creativity and labor in our digital economy. If we don't act now, we risk losing not only the integrity of journalism but also the trust that underpins public discourse.

Key Facts

  • AI training data scale: OpenAI's mid-training datasets contained over 91,000 copies of works from publishers like The New York Times and Daily News
  • Content scraping method: Companies scraped paywalled content through methods that circumvented existing paywalls and protections
  • Data source volume: A Common Crawl-derived dataset included over 2 million documents specifically from nytimes.com
  • Internal Microsoft admission: A top Microsoft executive privately described AI training practices as 'theft'
  • Publisher threat assessment: OpenAI's leadership recognized an 'existential threat' to publishers from their AI models
  • Market impact on journalism: Microsoft's Copilot caused click-through rates for The Times' content to drop as much as 93%
  • Labor theft characterization: Microsoft's Director of Applied Science Brent Hecht called AI scraping 'an astonishing theft of unprecedented proportions'
  • Copyright notice removal: Training data was deliberately stripped of copyright notices before being fed into models

Background

New court filings in The New York Times vs. OpenAI and Microsoft case reveal that both companies scraped paywalled content to train AI models, with Microsoft internally labeling the practice as theft. The revelations show companies went to great lengths to circumvent paywalls, build training datasets via mass scraping, and remove copyright notices from training data. These filings demonstrate significant market harm to publishers, particularly The New York Times, whose click-through rates dropped dramatically due to AI products like Microsoft's Copilot.

Quick Answers

What did Microsoft privately call AI training practices?
Microsoft privately labeled OpenAI's data practices as theft.
When were these court filings unredacted?
The court filings were unredacted in September 2026.
Who is the author of this article?
Rebecca Bellan is the author of this article.
What was the scale of content scraped by OpenAI?
OpenAI's mid-training datasets alone contained over 91,000 copies of works from publishers like The New York Times and Daily News.
How did Microsoft's Copilot affect The Times' content?
Microsoft's Copilot caused click-through rates for The Times' domain to drop as much as 93% compared to traditional Bing search.
What did OpenAI leadership say about publishers?
OpenAI's leadership recognized an 'existential threat' to publishers from their AI models.
Who described AI scraping as theft of unprecedented proportions?
Microsoft's Director of Applied Science Brent Hecht described AI scraping as 'an astonishing theft of unprecedented proportions'.
What happened to copyright notices in training data?
Copyright notices were deliberately stripped from training data before it reached the models.

Frequently Asked Questions

What was Microsoft's internal view of AI training practices?

Microsoft privately labeled OpenAI's data practices as theft, with a top executive describing the AI training process as 'theft.'

How much content did OpenAI scrape from The New York Times?

OpenAI's mid-training datasets contained over 91,000 copies of works published by The New York Times.

What impact did AI products have on publisher revenues?

Microsoft's own data shows that its Copilot 'answer engine' caused click-through rates for The New York Times' domain to drop as much as 93%.

Did OpenAI acknowledge threats to publishers?

OpenAI's leadership recognized an 'existential threat' to publishers from their AI models, according to the unredacted court filings.

What methods did companies use to circumvent paywalls?

Companies allegedly built training datasets that disproportionately relied on scraped news content and pulled millions of articles from Common Crawl, a free web crawl data repository.

How did AI companies obtain content without permission?

Companies scraped content from paywalled sources through methods that circumvented existing protections, including bypassing paywalls undetected and building datasets via mass scraping.

Source reference: https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/

Comments

Sign in to leave a comment

Sign In

Loading comments...

More from Business