Request a scoping call Contact
← Research

Where your training data came from is now a balance sheet question

Data acquisition, training and publication can raise different rights questions. Record corpus provenance and assess supplier continuity, contract protection and output use against the relevant jurisdiction and facts.

Data / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

AI copyright diligence needs specific facts rather than a general position on whether training is transformative. Identify the material, how it was obtained, which use is proposed and where the relevant acts occur. Those facts let legal advisers assess the corpus and contract.

Recent litigation illustrates why acquisition, retained libraries, training and outputs should be considered separately. The cases differ by claim, jurisdiction and evidence. They provide questions for diligence rather than a settled rule for every AI service.

Treat corpus provenance as a commercial and legal review input. The buyer needs evidence of acquisition and permitted use, alongside the model’s capabilities. A court decision about one activity cannot establish the lawfulness of all the others.

What the decisions have actually separated

The Bartz settlement website records final approval on 20 July 2026 of the USD 1.5 billion Anthropic settlement. That settlement should be distinguished from the earlier ruling separating training and the retained pirated-book library. In the UK, the November 2025 Getty judgment rejected the secondary-infringement claim concerning model weights while finding limited trademark infringement. It did not provide a general ruling that training is lawful. Norton Rose Fulbright’s 2026 survey offers further context across cases. Review the relevant orders and current proceedings for a particular legal decision.

What the litigation has and has not settled Fig. 01
The question people ask What the decisions are turning on
Is training on copyrighted work lawful? Courts are treating that separately from how the work was obtained
Do model weights contain the works? A UK court declined to treat weights as infringing copies
Is this settled now? It is fact-specific, jurisdiction-specific, and moving
Does this affect us? We only use a model Assess your actual use, alongside supplier continuity, indemnities and output rights

A model user may face supplier disruption, changed pricing, contractual uncertainty or output-related claims. Its exposure depends on its own activity and the agreement. Do not assume that using a model either transfers the provider’s liability or removes every rights issue.

The three exposures a buyer actually carries

Continuity. Assess what happens if the model is restricted, withdrawn or changed. Identify an evaluated alternative and the transition work it requires. A gateway may reduce interface changes, but does not alone preserve service behaviour or legal permission to use the replacement.

Indemnity. Read the applicable terms and exclusions. Coverage can depend on service tier, filters, customer inputs and permitted use. Obtain legal advice on the actual contract rather than relying on a supplier’s summary of copyright protection.

Output. Assess what the organisation generates, publishes or acts on. Review and permitted source material can reduce some risks but do not establish that every output is lawful. Preserve the relevant record and route consequential or disputed uses to an appropriate rights review.

The part that is your own data

For material the organisation supplies, maintain a record covering fine-tuning data, retrieval corpora and evaluation sets. A licence or processing contract may permit some uses and restrict others. The organisation can investigate that position directly, although rights can still be uncertain or disputed.

Review changes of use explicitly. Examples include a prototype corpus entering production, licensed records moving into training, or customer material being reused beyond its original processing purpose. These are possible diligence gaps, not findings about the frequency of unlawful repurposing.

A provenance register should identify the corpus, source, acquisition record, permitted uses, restrictions and responsible owner. Link supporting documents and record unresolved questions. Keep it current when material is added, removed or reused.

What this does not tell you

Nothing here is a legal opinion, none of it transfers between jurisdictions, and the position is moving quickly enough that a piece written now describes a moment rather than a settled rule. Anyone who needs to know whether a specific corpus is usable should be asking counsel, not an architecture practice.

The appropriate response depends on the intended use and available evidence. Some risks can be addressed through licences, contract terms, review or a narrower scope. Others may prevent the proposed use. The architecture record should preserve that decision without minimising or exaggerating the legal exposure.

The data owner should establish provenance before approving a new corpus use. Ask legal advisers to resolve material rights questions and record any accepted limitations. Use the register in supplier review and change control so the evidence remains available when the service evolves.

Filed under · Data · Data · Provenance · Supplier risk Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.