Request a scoping call Contact
← Research

Plan the embedding change as a migration

Changing an embedding model can require a new index, a retrieval evaluation and a controlled cutover. Record the source versions, transformations and permissions needed to rebuild the service before committing to the switch.

Data / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

An embedding model with stronger benchmark results or lower prices can be worth adopting. The investment decision also needs the cost of rebuilding the index, testing retrieval and retaining a recovery route. A configuration change can start a migration whose work extends well beyond that setting.

Check compatibility between the stored vectors and the proposed query model. Different embedding models generally produce different vector spaces, so comparing their outputs without a validated compatibility arrangement can make retrieval unreliable. Where a replacement requires new vectors, the team may be able to reuse retained chunks. Re-extraction and chunking are necessary only when those inputs need changing or cannot be recovered.

Treat a change of embedding model as a migration with explicit compatibility, evaluation and recovery conditions. The index is a derived artefact. Its design record should explain what produced it and what must be preserved to reproduce the relevant state.

What the index actually depends on

Inputs needed to explain and reproduce an index Fig. 01
  1. Input 01 The source documents At the version they were on when they were ingested — not the current version.
  2. Input 02 The extraction How text came out of the original format. A parser change can alter downstream text and chunks.
  3. Input 03 The chunking rules Size, overlap, and what counted as a boundary. Retain the versioned configuration.
  4. Input 04 The metadata attached Permissions, dates, source system. What the filters run against.
  5. Input 05 The embedding model Including the exact version. Check compatibility with the query model.

The embedding version is one of several inputs that affect retrieval. Extraction determines which text survives, chunking determines the available passages, and metadata governs filters and access. Record those inputs together. A model name alone cannot explain why a particular passage was available to a particular user.

The migration that is not planned as one

illustrative fixture

Test the index before cutover

Can the candidate index meet retrieval and access conditions with a recovery route?

Fixture
Known source revisions, transformation settings and a representative retrieval set.
Procedure
  1. Check stored-vector and query-model compatibility.
  2. Reuse retained chunks where valid, or rebuild missing transformations.
  3. Compare the candidate with the current index on reviewed retrieval and access cases.
  4. Test cutover and the agreed route back to the retained index.
Result to check
Record measured retrieval changes and unresolved access or reconstruction gaps.
Evidence limit
These are proposed checks, not an observed migration or a claim that every replacement needs re-extraction.
Next test
Resolve gaps and obtain acceptance from the retrieval and data owners.

Representative method for this article, not a measured deployment result.

Can the candidate index meet retrieval and access conditions with a recovery route? Evidence limit: These are proposed checks, not an observed migration or a claim that every replacement needs re-extraction.

Reviewed 2026-10-02

Evaluate the proposed index before cutover. A model with stronger published benchmark results can behave differently on a corpus containing specialist vocabulary, tables or short documents. A test that checks only whether the application returns an answer leaves that change in retrieval quality unmeasured.

Build a retrieval evaluation set from representative questions and reviewed relevance judgements. Its size and coverage should follow the consequences of a missed or incorrectly returned document, rather than a universal question count. Evaluate retrieval separately from the final answer: fluent answers can conceal missing evidence, while a suitable retrieved passage can still be misused by the generation step.

Why this is a governance problem, not only an engineering one

If a system’s answers can be questioned later — and any system supporting a consequential decision can be — then the question is what the index contained at the time of the answer. That has an answer only if the index carries a version that changes when any of the five inputs above changes, and only if the trace recorded which version answered.

Without it, the honest response to a question about a past answer is that the system can be re-run against a corpus that no longer exists. That is not a reconstruction. It is a new experiment presented as an audit.

Compare migration cost with the expected benefit

Estimate migration costs against the actual corpus, evaluation work and cutover plan. Check the support policy for the specific embedding provider and model. Published model deprecation calendars also matter elsewhere in an AI service, but the linked Claude calendar concerns language models. It does not establish an embedding-model retirement schedule.

Keep extraction and chunking configurations with the derivation record, and retain source versions according to the service’s evidence and data-protection requirements. Include switching costs in provider selection. These decisions make a later migration easier to scope and reveal which parts still depend on unavailable code or material.

What this does not tell you

There is no general answer to which embedding model is right, and this is not an argument for staying on an old one. Retrieval quality is corpus-specific to a degree that surprises people, and the only reliable way to choose is to measure on your own material.

The required reconstruction capability depends on the use. A disposable prototype may tolerate a rebuild that changes its results, while a consequential decision-support service may need a traceable historical state. Record the accepted limitation and revisit it before extending the prototype into a production service.

The retrieval owner should ask the team to demonstrate a rebuild from retained inputs. Use that exercise to identify missing versions, permissions or configuration, then decide which gaps must close before a provider switch. A measured migration and a defined rollback route provide a firmer basis for approval than the replacement model’s benchmark alone.

Filed under · Data · Retrieval · Data · Provenance Inference Institute · 02 Oct 2026 (updated)

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.