Start a conversation Contact
← Research

An erasure request must reach the derived copy

Deleting a source record may leave personal data in a training extract, retrieval index or saved AI output. Map those copies to the person and the purpose before deciding what a rights request requires.

A customer asks a firm to erase personal data. The account team removes the record from the application database and closes the request. Weeks later, an AI-assisted service retrieves a passage from an old index that still contains the same details. The database deletion succeeded. The organisation’s account of where it processed the person’s data did not.

The answer is to map the copies and transformations that matter to the request. A training extract, embedding store, retrieval cache, profile, human-review queue and saved model output may each have different ownership and retention. An erasure workflow that sees only the original table cannot decide what else needs to be examined.

Follow the person, not the file name

The ICO’s AI and individual-rights guidance distinguishes personal data in training data, outputs and, in some cases, the model itself. The right to erasure is not absolute. The ICO also says that erasing an individual’s training data does not automatically require deleting every model trained on it, unless the model contains the data or it can be inferred from the model. The duty is to assess the actual processing and any applicable exception, not to choose between deleting everything and deleting only the source record.

That distinction makes data lineage operational. A derived representation may no longer look like the original row, yet still link to a person or reproduce their information. Conversely, a statistical model need not contain a recoverable copy of every training example. The team needs to know which artefacts remain personal data in its context and how a request can identify them.

Begin with the identifiers carried forward. Does a training extract retain an account key? Does a retrieval index retain document IDs that can be joined to the customer? Can a saved prediction be found through the profile to which it was attached? If a pipeline strips direct identifiers, record how it was tested for indirect identification and what additional information a requester could supply. The ICO notes that an inability to identify somebody in training data can change when the person provides information that makes identification possible.

Design a request as a traversal

Write a data map from source to extract, transformation, index, model, prediction and downstream action. For each node, record the purpose, owner, retention rule, identifier or lookup path, and whether deletion can be propagated. Where a supplier processes a node, record how it assists the controller. The ICO says a controller’s choice of a service that makes rights requests difficult does not remove its obligations.

Run a representative request through the map. The result should be a list of places checked, actions taken, exceptions considered and residual artefacts requiring judgment. An index may be rebuilt from a cleaned source, but the current index might need direct removal first. A cached answer may have a separate lifetime. A model that demonstrably contains a person’s data may raise a different remedy from one that does not. Those are decisions to record, not assumptions to bury in a deletion script.

Take a customer whose support conversation was copied into a search index and later quoted in a saved assistant answer. The request handler should start from the account record, follow the document identifier into the index and look for stored answers that refer to that document. Removing the source conversation does not remove a cached passage. Rebuilding the index next month does not settle whether the current index can still return it today.

The closure record should name the copy, action, owner and date for each location. It should also say where the organisation concluded that an item was outside scope or that an exception applied, with the reason available to the privacy lead. This is more useful than a single “deleted” flag because it shows which branches of the pipeline were actually visited. A processor’s confirmation should identify the affected service and version, not merely assert that it follows a deletion policy.

Test the map after the next pipeline change. A new cache or evaluation store can create a copy that the original request form cannot reach. The owner of the data flow should update the traversal whenever a component begins to store or derive personal information. That makes future rights requests a release consideration rather than an emergency archaeology exercise.

The earlier article on re-embedding as a migration asks how to preserve retrieval quality and version boundaries when vectors change. A rights request asks a different question: which person can still be represented or retrieved, and what processing must stop? The same inventory can support both, but the acceptance test is different.

The limit of the map

A lineage diagram does not decide whether erasure applies to a particular record, model or output. Identifiability, lawful basis, exceptions and the person’s actual request need legal and factual assessment. This article does not promise that a single technical deletion will settle those questions.

The data owner and privacy lead should nevertheless be able to run one request end to end and show where it travelled. If they cannot name the derived copies, they cannot explain why the request stopped where it did.

Filed under · Data · Data protection · Erasure · AI data lifecycle Inference Institute · 25 Sept 2026

Related engagement

The decision behind this article

A structured assessment of one AI system covering risks, impacts, controls and residual risk.

Explore AI Risk & Impact Assessment →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.