Start a conversation Contact
← Research

Where your training data came from is now a balance sheet question

Courts have begun separating the act of training from the act of acquiring the material, and the money has landed on acquisition. That distinction moves the diligence question from what a model does to where its inputs were obtained.

For two years the argument about AI and copyright was conducted at a level of abstraction that made it useless to anybody buying a system. Either training was transformative and therefore fine, or it was wholesale copying and therefore not, and which of those you believed correlated closely with what you did for a living.

The decisions arriving now are more specific and less comfortable than either position, because they turn on facts about acquisition rather than on the nature of learning. That is a considerably more answerable question, and it is one a buyer can ask.

The claim: provenance of training material has become a due diligence item with a price attached, and the price is being set by how material was obtained rather than by what was done with it afterwards.

What the decisions have actually separated

In the United States, the Anthropic authors litigation produced a settlement of $1.5 billion, approved in 2026, in a case where the court had earlier indicated that training could be transformative while the retention of a library of pirated books was a separate matter. In the United Kingdom, Getty’s case against Stability AI failed on the secondary infringement argument — the High Court declined to treat model weights as an infringing copy — while a narrow trademark point succeeded. Norton Rose Fulbright’s survey of where the copyright cases stand in 2026 is a reasonable single place to see the pattern across jurisdictions.

What the litigation has and has not settled Fig. 01
The question people ask What the decisions are turning on
Is training on copyrighted work lawful? Courts are treating that separately from how the work was obtained
Do model weights contain the works? A UK court declined to treat weights as infringing copies
Is this settled now? It is fact-specific, jurisdiction-specific, and moving
Does this affect us? We only use a model Your exposure runs through indemnities, continuity and output, not through training

The fourth row is the one that matters to almost everybody reading this. Very few organisations train foundation models. A great many depend on one, and their exposure is not that they will be sued for training. It is that a provider might be, with consequences that reach them through service continuity, price, or the terms of an indemnity nobody has read closely.

The three exposures a buyer actually carries

Continuity. If a model is withdrawn, restricted or altered as a result of litigation or settlement, what happens to the systems built on it? This is the same question as model retirement, arriving through a different door, and it has the same answer: a gateway, an evaluation set, and a tested alternative.

Indemnity. Most major providers offer some form of copyright indemnity for outputs. They differ substantially in what they cover, what conditions attach, and what is excluded — typically including cases where the customer supplied the infringing input or disabled a filter. The document is worth reading rather than summarising, and the summary a sales team gives is not the document.

Output. The exposure that is genuinely yours rather than inherited is what your system produces and what you do with it. A model that reproduces protected expression is a risk you are running when you publish the output, and the controls are ordinary ones: review before publication, retrieval from material you have rights to, and a record of what was generated.

The part that is your own data

The version of this question that organisations control entirely is the material they supply themselves — fine-tuning corpora, retrieval indexes, evaluation sets. The rights position for that material is knowable, and it is frequently not known.

Licensed data acquired for one purpose is regularly repurposed as training or retrieval material, and the licence often does not permit it. Content scraped for a prototype becomes a production index. Customer material processed under a contract that says nothing about model training becomes a fine-tuning set, because it was the data that was available.

Each of those is a decision somebody made quickly, and the record of it is usually a commit message. A provenance register — what material is in each corpus, where it came from, and under what right it is being used — is the document that makes this answerable. It is not difficult. It is only difficult to reconstruct afterwards.

What this does not tell you

Nothing here is a legal opinion, none of it transfers between jurisdictions, and the position is moving quickly enough that a piece written now describes a moment rather than a settled rule. Anyone who needs to know whether a specific corpus is usable should be asking counsel, not an architecture practice.

Nor is the argument that generative systems carry unmanageable legal risk. For the overwhelming majority of enterprise uses they do not, and the practical exposure is smaller than the coverage suggests. What has changed is that provenance is now a question with commercial weight behind it, which means it belongs in supplier diligence and in your own data register rather than in a paragraph about ethics.

The reader who should act is whoever owns the data supply for the next model project. Write the provenance register before the corpus is assembled. It is a spreadsheet, it takes a day, and it is the artefact that determines whether the question can be answered at all when somebody eventually asks it.

Filed under · Data · Data · Provenance · Supplier risk Inference Institute · 13 Aug 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.