Policy

Seattle Times and Newsday Sue OpenAI and Microsoft for Copyright Infringement

Two major regional newspapers have filed copyright infringement suits against OpenAI and Microsoft over unauthorized use of their journalism to train AI models.


Seattle Times and Newsday Sue OpenAI and Microsoft for Copyright Infringement

Two major regional American newspapers — The Seattle Times and Newsday — have filed copyright infringement lawsuits against OpenAI and Microsoft, adding to a growing body of litigation targeting the use of published journalism to train large language models. The suits follow a pattern established by earlier cases, most prominently The New York Times' lawsuit filed in late 2023, and signal that legal pressure from the media industry is broadening rather than receding.

The core allegation is consistent across this wave of publisher litigation: that OpenAI ingested decades of copyrighted news content without license or compensation to build its models, and that Microsoft, as a primary investor and distribution partner, bears shared liability. For regional publishers operating on tighter margins than national outlets, the stakes are both financial and existential — their archives represent one of the few monetizable assets they retain.

The legal theory rests on whether training data ingestion constitutes a transformative use protected under fair use doctrine, or whether it constitutes reproduction at scale that requires licensing. Courts have not yet issued definitive rulings on this question, and these cases will likely remain in pretrial phases for years. The outcome, however, will set binding precedent for how AI developers can access publicly available content going forward.

For OpenAI and Microsoft, the accumulation of publisher lawsuits creates compounding legal exposure. Each new filing adds to discovery obligations, potential damages calculations, and the reputational cost of being positioned as an extractive actor relative to an already-struggling journalism sector. OpenAI has pursued licensing agreements with some publishers — including The Associated Press and certain News Corp properties — as a parallel strategy, but voluntary licensing deals do not neutralize claims from publishers who never agreed to participate.

The business implications extend beyond the defendants. Any judicial ruling that training data ingestion requires prior licensing would structurally alter the economics of foundation model development. Models trained predominantly on web-scraped data would face retroactive liability, and future data acquisition would require a licensing infrastructure that currently does not exist at the scale AI development demands. This would disproportionately affect smaller AI developers who lack the capital to negotiate content agreements with major publishers.

For enterprises already deploying AI systems built on these models, the litigation does not create immediate operational risk. Liability, if established, would sit with the model developers, not with downstream users accessing APIs or software products. However, prolonged legal uncertainty could slow the expansion of AI capabilities into journalism, content summarization, and media monitoring — domains where the boundary between training and inference becomes harder to distinguish.

The AIRA perspective here is structural: this litigation cluster is functioning as an informal regulatory mechanism in the absence of formal AI training data law. Congress has not moved to define training data rights, the Copyright Office has issued guidance that stops short of resolving the core question, and so the judiciary is being asked to adjudicate a technology policy question through the lens of existing intellectual property statute. The outcomes will be inconsistent, jurisdiction-dependent, and slow — which is itself a form of friction that shapes how AI companies invest in data acquisition strategies.

What these suits collectively signal is that the period of unrestricted model training on publicly available content is narrowing. Whether through court order, settlement-driven licensing norms, or eventual legislation, the data acquisition layer of AI development is moving toward a more contested and negotiated space. Organizations building long-term AI infrastructure should treat training data provenance as a material risk category, not a legal footnote.

Sources: — The Verge (https://www.theverge.com/ai-artificial-intelligence/990932/seattle-times-newsday-lawsuit-openai-microsoft)