AI agents · OpenClaw · self-hosting · automation

Quick Answer

Is AI Training Fair Use? Thomson Reuters v. Ross Explained

Published:

The short answer

On September 29, 2026 the Third Circuit became the first US appeals court to rule on fair use for AI training — and it ruled against the AI company. Ross Intelligence’s use of 2,243 Westlaw headnotes to train a competing legal-research engine was not fair use. But the court anchored the result in Ross’s specific facts (direct competitor, same function, free alternatives available, licensing market emerging) and expressly declined to extend it to generative AI, citing Bartz v. Anthropic, Kadrey v. Meta and the DOJ’s position in the OpenAI litigation. Facts verified October 2, 2026.

The case in one table

Detail
CaseThomson Reuters Enterprise Centre GmbH & West Publishing Corp. v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)
Filed in appeals courtArgued June 11, 2026; opinion filed September 29, 2026 (initially sealed for redaction review)
BelowJudge Stephanos Bibas, D. Del., partial summary judgment February 2025
Issues decided(1) 2,243 non-verbatim Westlaw headnotes are copyrightable; (2) Ross’s copying for ML training was not fair use
Issues not decidedCopyrightability of headnotes quoting opinions verbatim; originality of the Key Number System (forfeited)
PostureInterlocutory appeal from partial summary judgment — the litigation is not over
Generative AI?No. Ross’s system returned existing judicial passages; it did not generate text
Ross todayShut down its platform in 2021 under the weight of the suit

What Ross did

Thomson Reuters and West sued Ross in Delaware on May 6, 2020. Ross had built an AI legal-research tool: ask a natural-language legal question, get back passages from roughly ten million judicial opinions. To train it, Ross’s contractor LegalEase and a subcontractor produced about 25,000 question-and-answer memoranda. Westlaw’s headnotes — the short editorial summaries of a case’s legal points — were used to frame the questions, and judicial passages tied to those headnotes were labelled as highly, partially or not responsive. Ross did not contest on appeal that the copying happened or that it was attributable to Ross.

Procedurally, the result evolved. In September 2023 Judge Bibas left the core questions for trial. In February 2025, after reconsidering, he revised his ruling and granted Thomson Reuters partial summary judgment on both copyrightability and fair use, writing that “Ross took the headnotes to make it easier to develop a competing legal research tool. So Ross’s use is not transformative.” The Third Circuit has now affirmed.

The panel described the dispute as no more than an ordinary copyright case despite the AI component, and the four fair-use factors came out as follows:

  1. Purpose and character. Not transformative. Ross used the headnotes for the same ultimate purpose Westlaw did — locating responsive legal authority — and did so commercially, as a direct competitor. The court rejected the idea that “training” is automatically transformative because machine learning is involved.
  2. Nature of the work. Headnotes are editorial but modestly creative; this factor did little work.
  3. Amount used. Ross copied entire headnotes — short works copied whole — which weighed against fair use.
  4. Market effect. Two markets counted: the legal-research market Ross competed in, and a developing market for licensing editorial material as AI training data. The second holding is the one future plaintiffs will quote.

Ross also had a cheaper lawful path: the underlying judicial opinions are public domain. Choosing protected expression when unprotected material would have served the technical purpose cut against it.

What the ruling does not say

The opinion is a major but not total win for Thomson Reuters, and it is not a rule that AI training infringes:

  • It is an interlocutory ruling on partial summary judgment; portions of the case remain in the district court.
  • It leaves open whether verbatim-quote headnotes are copyrightable and does not reach the Key Number System.
  • In a footnote the court distinguished generative AI. It discussed Bartz v. Anthropic and Kadrey v. Meta — the 2025 Northern District of California decisions that found training on lawfully acquired books transformative — and the Department of Justice’s September 2026 filing in In re OpenAI, which argues generative-model training may be transformative and should be analysed separately from any infringing outputs. The panel stressed that Ross’s system could not create new expression, whereas generative systems are argued to serve a materially different purpose. It declined to transpose its result to LLMs.

Commentators including LawSites’ Bob Ambrogi cautioned that the facts — pre-generative AI, a non-generative system, head-to-head competition — may confine the decision. Whether it reverberates depends on how later courts read the footnote.

What changes anyway

Even confined to its facts, Ross changes the generative-AI docket by handing plaintiffs appellate authority for four propositions:

PropositionWhy it matters for LLM cases
An intermediate training copy can infringe even if the protected text never appears in outputsUndercuts “the model doesn’t contain the book” defences
Copying an entire short work weighs heavily against fair useHeadnotes, lyrics, poems, news summaries, code snippets
Direct commercial substitution is decisive after WarholProducts that compete with the source — news chatbots vs publishers, legal AI vs Westlaw
An emerging AI-training licensing market is a cognisable factor-four marketThe ~100 publishers Google is now paying, the Anthropic and OpenAI licensing deals, all become evidence that a market exists

The practical reading from legal analysts: courts are stopping treating “AI training” as one category. The questions are now what was copied, how it was acquired, why that expression was necessary, what the model does with it, whether outputs or the service substitute for the source, and whether a licensing market exists. Ross is strong precedent because on every one of those dimensions the facts favoured Thomson Reuters.

For developers: the risk ladder after Ross

Highest risk: a competitor’s proprietary, human-curated data → used to reproduce the same commercial function → public-domain or licensed substitutes ignored → no record of why protected expression was technically necessary.

Lower risk: public-domain or openly licensed corpora; a genuinely different downstream purpose; lawful acquisition (Bartz turned on pirated copies); anti-memorisation measures; provenance records per source; compliance with machine-readable rights reservations such as robots.txt and the EU’s TDM opt-out; negotiated licences for high-value proprietary data. Legal-AI specifically — Harvey, CoCounsel, Claude for Legal and Astra for Law — now operates under a circuit precedent written about its own market.

The same week, Judge Mehta dismissed publishers’ antitrust suits over Google AI Overviews, telling them antitrust is not the remedy — see what the Chegg and Penske dismissal means. Copyright, after Ross, is the remedy publishers are left with.

Last verified: October 2, 2026, against the Third Circuit opinion as published and contemporaneous legal reporting.

Sources