
Seattle Times and Newsday Sue OpenAI and Microsoft Over Paywalled Content Training
Published by AINave Editorial • Reviewed by Ramit
The Seattle Times and Newsday have filed a copyright infringement lawsuit against OpenAI and Microsoft in the Southern District of New York, alleging that the companies scraped paywalled content to train ChatGPT, Microsoft Copilot, and the AI-powered Bing. The publishers claim the chatbots can reproduce entire passages and closely paraphrase reporters' work, and they are demanding that the tech companies destroy any copies of the material, as well as training datasets or AI models that incorporate it.
This lawsuit is the latest in a series of similar actions from news organizations, including the New York Times, Alden Global Capital's MediaNews Group, and Mother Jones' publisher. For AI builders, the pattern is becoming hard to ignore: the legal status of training on publicly accessible but copyrighted content remains unresolved, and the risk of retroactive liability is real.
What the Publishers Allege
The suit specifically claims that OpenAI and Microsoft bypassed paywalls to collect articles for training data. The publishers argue that this is not fair use because the models compete with their original journalism by summarizing and reproducing it without permission. OpenAI has responded that it trains only on publicly available materials grounded in fair use. Microsoft said it was surprised by the lawsuit but is willing to explore solutions.
This distinction between publicly available and paywalled content is central. Even if public web scraping is fair use, actively circumventing paywalls changes the legal calculus. Builders should note that the technical boundary between open scraping and unauthorized access can be narrow and is increasingly being tested in court.
Why AI Builders Should Pay Attention
If the publishers win, the consequences could extend beyond OpenAI and Microsoft. The demand to destroy training datasets and models sets a precedent that could affect any company that trains on copyrighted content without explicit licenses. Even if your model is fine-tuned on a base model that was trained on potentially infringing data, downstream liability is not off the table.
More immediately, this case reinforces the need for clear data provenance. If you are building a product that relies on RAG, web scraping, or training on news content, you should review your data sources and consider licensing agreements. Some publishers are already signing deals with AI companies; others are suing. The landscape is fragmented, and the safe path is not always obvious.
What the Lawsuit Seeks
The plaintiffs are asking for:
- Destruction of all copies of the copyrighted material in OpenAI's and Microsoft's possession.
- Destruction of training datasets and AI models that incorporate the material.
- Damages for alleged infringement.
This is not just a request for financial compensation. If granted, it could force OpenAI and Microsoft to retrain or alter models that have already been deployed. For any team using these models via API, that could mean service disruptions or changes in model behavior.
Caveats: What's Still Unclear
The outcome is far from certain. The fair use defense is strong, and the Trump administration recently filed a Statement of Interest in the New York Times case arguing that a win for publishers could harm national security and local journalism. That filing signals that the government may intervene in this case as well.
Additionally, many news organizations have chosen to license their content rather than litigate. OpenAI has signed deals with publishers like Axel Springer, Le Monde, and the Financial Times. This dual approach -- suing some while licensing others -- creates uncertainty for builders who need predictable access to training data.
For now, the safest assumption is that training on any copyrighted content, especially behind a paywall, carries legal risk. If you are building an AI product that relies on news content, consider whether your use case falls under fair use or whether you need a license. This case will not resolve quickly, but it will shape the rules for years to come.
FAQs
Sources
- OpenAI, Microsoft face copyright lawsuit from local newspapers
- Local newspapers sue OpenAI and Microsoft over use of... | Mashable
- OpenAI, Microsoft Face Lawsuit by Mother Jones Publisher CIR
- Two More News Organizations Sue OpenAI And Microsoft For...
- NYT sues OpenAI, Microsoft for alleged copyright infringement
- Court Filings In A.I. Suit Invoke Copyright Law, Culture and Sports
- Seattle Times and Newsday sue OpenAI and Microsoft for copyright infringement
- Daily News, NY Times want ‘fair use’ argument rejected in copyright case against OpenAI, Microsoft
- Seattle Times and Newsday Claim OpenAI Scraped Paywalled Articles to Train ChatGPT
- Two more newsrooms join the case against OpenAI and Microsoft
- 8 US newspapers sue ChatGPT-maker OpenAI and Microsoft for copyright infringement
- PYMNTS | OpenAI, Microsoft Face Lawsuit by Mother Jones...
- OpenAI, Microsoft face fresh lawsuit over copyright infringement...
- Trump Administration Sides With OpenAI in New York Times Copyright Lawsuit
- US government backs OpenAI in New York Times copyright case
- Newsday sues OpenAI and Microsoft over copyright infringement





















