Author: IRPA AI Senior Analyst, Kieran Gilmurray 

The creative economy and the AI sector are in a defining legal standoff. Since 2023, painters, authors, songwriters, and coders have launched lawsuits claiming that generative models were built on unlicensed creative labour. The core question is unprecedented at this scale: can developers copy entire works to train models, and if so, on what terms?


TLDR / At a Glance

  • Dozens of lawsuits by artists, authors, publishers, musicians, and developers challenge the unlicensed use of copyrighted works in AI training.
  • U.S. courts are split, with some rulings favouring creators and others hinting at fair use for transformative training.
  • The EU allows text and data mining with rights-holder opt-outs, and the AI Act will require disclosure of copyrighted sources.
  • The industry is shifting toward licensing, provenance technology, contributor funds, and model guardrails.

Introduction: A Clash Between Innovation and Ownership

The creative economy and the AI sector are in a defining legal standoff. Since 2023, painters, authors, songwriters, and coders have launched lawsuits claiming that generative models were built on unlicensed creative labour. The core question is unprecedented at this scale: can developers copy entire works to train models, and if so, on what terms?

AI firms argue that training is a transformative, intermediate use that benefits the public. Creators counter that it is mass appropriation, enabling tools that compete with human creativity without consent or compensation. Courts, regulators, and the market are now shaping the answer.

The Lawsuits That Redefined the Landscape

Visual Artists vs. Image Generators

In Andersen v. Stability AI, artists allege that billions of images, including their works, were copied without consent to train Stable Diffusion and related systems. A California judge let core copyright claims proceed, citing plausible evidence that training compresses and can reconstruct protected images. Trial is set for 2026.

Getty Images v. Stability AI in the UK and U.S. alleges unlicensed scraping of millions of photos, trademark misuse of the Getty watermark, and database rights violations. While jurisdictional hurdles narrowed some claims in London, U.S. litigation continues. These cases test whether reproducing large datasets to teach a model is infringement.

Authors and News Outlets vs. AI Text Models

In Authors Guild v. OpenAI, novelists contend that GPT was trained on pirated e-book archives and can generate text in the style of their works. OpenAI asserts fair use, arguing training is analytical and outputs are not substitutes. Meanwhile, The New York Times v. OpenAI alleges copying of articles and removal of copyright management information. While litigating, OpenAI has signed licensing deals with AP, Condé Nast, Axel Springer, and others, signalling a shift toward paid access.

Music and Code

Music publishers sued Anthropic, alleging that Claude reproduces song lyrics verbatim on request. The case targets both training and output. In software, Doe v. GitHub challenged Copilot’s training on licensed code and alleged license stripping. Most copyright and DMCA claims were dismissed for lack of specific verbatim copying, though contract and license claims remain under appeal.

Together, these suits argue that models built on unlicensed content erode existing markets and bypass the creators who supplied the raw material.

The Legal Fault Lines: Fair Use and Authorship

The central issue is whether training on copyrighted data is fair use. Developers analogise to earlier cases, allowing Google to scan books for indexing or cache images for search. Creators reply that wholesale ingestion of entire works at a commercial scale is different, particularly where outputs can serve as substitutes or displace licensing markets.

Recent rulings cut both ways. In Thomson Reuters v. ROSS, a Delaware court rejected fair use, finding the AI legal tool competed with Westlaw’s core market and used expressive summaries non-transformatively. In Bartz v. Anthropic, a California court preliminarily indicated that training large language models on books can be fair use where the purpose is analytical, the books were lawfully acquired, and there is no clear market substitution, though the case later settled. In Andersen, fair use was left for later, but the court credited allegations that training and outputs may plausibly infringe.

Authorship is clearer; purely AI-generated content without significant human creative input is not copyrightable in the U.S., and similar principles apply in the EU and UK. That means model outputs are largely unprotectable unless human authorship is sufficient, an irony that affects claims of ownership in outputs but not the legality of training inputs.

The Regulatory Divide: U.S. Fair Use vs. EU Opt Outs

The U.S. Copyright Office has reaffirmed that only human-authored portions of works can be registered, launched an AI policy initiative, but has not declared training fair use or infringement. Congress is considering bills that would mandate the disclosure of training data and clarify rights.

The EU set the table earlier. The 2019 DSM Directive allows text and data mining for commercial use unless rights-holders opt out via machine-readable signals. Courts confirmed that generic terms of service are not enough; technical opt-outs are required. The forthcoming AI Act will require foundation model providers to summarise copyrighted data sources used and confirm compliance with opt-outs, even if training occurred outside the EU, when models are offered in the EU.

Elsewhere, Japan permits broad data mining, the UK reversed a proposal for unrestricted commercial TDM after creator pushback, and global norms are diverging.

Industry’s Response: From Scraping to Licensing

Facing litigation and reputational risk, AI companies are moving toward licensing and provenance.

Licensing deals. OpenAI has inked agreements with news publishers, Shutterstock licenses its image library and pays contributors from a fund, Getty partnered with NVIDIA on a fully licensed generator, and Adobe trained Firefly on licensed or public-domain content while indemnifying enterprise customers.

Provenance and opt-outs. IPTC added “Data Mining” fields to image metadata so creators can flag permissions. The C2PA coalition is standardising cryptographic content credentials that record origins and edits. Crawlers like GPTBot claim to honour robots’ directives, and firms are piloting watermarking for AI outputs.

Guardrails and dataset curation. Major models refuse to output entire books or full lyrics, restrict prompts naming living artists who opt out, and retrain on filtered, licensed datasets. Policies forbid using the tools to infringe.

The market signal is clear, uncompensated scraping is giving way to paid access, attribution, and safety.

Creators’ Demands: Consent, Credit, Compensation

Creator coalitions advocate three principles. Consent, opt-in for training, or at least robust opt-outs. Credit, transparency about training sources and clearer attribution. Compensation, collective licensing or revenue sharing when works are used.

Proposals include collective licensing schemes, contributor funds, and protections for rights of publicity in AI voice and likeness cloning. Early settlements, like the Anthropic author class payout combined with opt-outs, and contributor programs at Adobe and Shutterstock, point to evolving norms. Unions have already secured contract limits on AI substitution in Hollywood.

Outlook: Toward a New Creative Contract

By late 2025, most cases are pending, appellate guidance is forthcoming, and regulation is tightening. The trajectory resembles earlier tech disruptions, early unlicensed use, then regulatory and market realignment toward licensing and revenue sharing.

A sustainable model likely blends lawful data access, provenance and transparency, safe outputs, and fair compensation. Innovation and creativity are not mutually exclusive, but harmony requires consent, clarity, and credible governance.

As one practitioner noted, AI’s long-term legitimacy will depend on how fairly it learns, not only on what it creates.


About the Author: Kieran Gilmurray


Senior IRPA AI Analyst & Advisor, Kieran Gilmurray is a certified executive coach, intelligent automation & digital transformation thought leader & content guru focused on helping solve complicated problems others can't.  For the past 25+ years, he has driven business digital transformation programs across a range of industries like digital technologies, intelligent automation, data analytics, social media and robotic process automation, having generated millions of dollars of value.

CLICK HERE TO SCHEDULE AN ANALYST BRIEFING 

Links:


Originally posted on 2025-11-07 in the IRPA AI Network — Announcements & Updates