AI training data copyright: Seattle Times and Newsday sue OpenAI and Microsoft
straitstimes.com

AI training data copyright: Seattle Times and Newsday sue OpenAI and Microsoft

Tech News
3 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRThe Seattle Times and Newsday sued OpenAI and Microsoft on September 4, alleging unauthorized scraping of paywalled content for AI training, highlighting legal risks for builders around data provenance.

Two major newspapers have escalated the legal fight over AI training data copyright. On September 4, 2026, The Seattle Times and Newsday filed a lawsuit in the Southern District of New York against OpenAI and Microsoft, alleging that the companies scraped their websites, including content behind paywalls, and used the articles to train AI systems such as ChatGPT, Microsoft Copilot, and Bing AI features.

The plaintiffs claim that OpenAI and Microsoft's AI products reproduce passages from their reporting, closely paraphrase articles, and provide users with answers that reduce the need to visit the newspapers' websites or buy subscriptions. The lawsuit seeks an order requiring destruction of copies of the works as well as training datasets or AI models that incorporate them.

OpenAI responded that its models are trained on publicly available data and grounded in fair use. Microsoft said it was surprised by the lawsuit but remains open to discussing solutions, emphasizing support for local journalism. The case echoes the New York Times copyright lawsuit filed against the same defendants in 2023, which recently reached the summary-judgment stage.

Why this matters for AI builders

This lawsuit signals that content owners, especially publishers with paywalled material, are increasingly willing to challenge how AI training data is sourced. For builders, the risk isn't just about future litigation. The outcome could influence whether certain publicly available but copyrighted data remains safe to use without explicit licensing. If the court rules against the fair-use defense, it could reshape the economics of training data acquisition, pushing more teams toward licensed datasets or synthetic data.

Practical implications for data sourcing and licensing

Builders should expect more scrutiny on data provenance, especially when scraping or using web-crawled corpora. Content licensing agreements may become more common for high-value publishers. Tools that verify data lineage and detect copyrighted material in training sets could see increased demand. Teams working on RAG pipelines that pull from news sources should also monitor how API providers handle copyright filters after this case.

What remains uncertain

The lawsuit is in its early stages. No decision on fair use or remedies has been made. The defendants have defended their practices as lawful, and similar cases are still pending. Until a court ruling provides clearer guidelines, the practical impact on most AI workflows is limited, though the legal uncertainty itself may affect how risk-averse teams select training data.

The case will likely take months or years to resolve, but for builders, it's a reminder that training data copyright is not just a legal issue; it's a procurement and compliance issue that may affect model performance, cost, and deployment timelines.

FAQs

They allege that OpenAI and Microsoft scrapped their websites, including paywalled content, and used the articles without permission to train AI systems and operate products such as ChatGPT, Microsoft Copilot, and Bing AI features. The plaintiffs claim the AI outputs reproduce or closely paraphrase their reporting, reducing the incentive for users to visit their sites or buy subscriptions.

Sources

Latest Tech News