Request a scoping call Contact
← Research

What an inference architecture has to hold

An inference design needs explicit choices about access, routing, context, serving and records. Assess those concerns against the workload so the service can be operated, evaluated and changed with known limits.

Architecture / Conceptual study
Make the connections explicit.
  1. Applications
  2. Interfaces
  3. Records

Trace the interfaces between applications and records before changing a system.

An architecture diagram that shows the application, retrieval store and model endpoint can explain the main flow while leaving operating questions open. The review also needs access rules, model selection, evidence assembly and retained records. These determine what the service can demonstrate when its use or dependencies change.

Adding those concerns later is possible, but can require changes across a live service. Historical information never recorded may remain unavailable. Establish the required boundary and accepted gaps before the design becomes a production dependency.

Review five concerns separately: gateway controls, routing, context assembly, serving and records. They need not become five products or five physical layers. Each needs an explicit design decision appropriate to the task.

The five layers

What sits between an application and a model, and what each part owns Fig. 01
  1. Layer 01 Gateway Identity, quotas, redaction, refusal policy, and a shared point for enforcing request controls.
  2. Layer 02 Routing Which model serves which request, what happens when it is unavailable, and the interface and evaluation work needed for a provider change.
  3. Layer 03 Context assembly Retrieval, entitlement filtering, ordering and truncation. What the model is actually given.
  4. Layer 04 Serving The hosted endpoint or your own accelerators. Batching, caching and the concurrency the system can absorb.
  5. Layer 05 Record What was asked, what was retrieved, which versions answered, what came back, and what the person did with it.

A demonstration proves only the paths it exercises. Before production, ask how the remaining access, operating and evidence requirements are met. Record omissions and the restrictions that make them acceptable for the proposed scope.

Assess the shared enforcement point

A gateway can centralise shared controls for applications that use it. Its effectiveness depends on coverage, enforcement and whether other paths bypass it. It is one design option for consistent policy, rather than the only possible place a rule can be applied.

Evaluate calling identity, quotas, data handling and recording at the relevant boundaries. A central enforcement point may reduce duplication, but receiving tools and source systems still need their own checks. Redaction and a gateway policy do not establish that no sensitive information can reach a provider.

A gateway may reduce some provider-migration work. OpenAI’s deprecation list illustrates why hosted-version replacement should be planned. The replacement still needs evaluation of model behaviour, interfaces and context limits. Do not assume the change is only a routing rule.

Decide the historical evidence requirement

Retained records support investigation of past behaviour. Specify the evidence needed for consequential outputs and the lawful retention period. Future-looking monitoring and incident response also matter, so historical reconstruction is one requirement among several.

For the agreed scope, connect the request, permitted evidence and revisions, relevant context, model and prompt versions, output and subsequent action. Protect the records and establish access for authorised reviewers. The required detail and retention depend on sensitivity, consequence and the questions the service must answer.

Recording can be improved for future requests, but missing historical inputs cannot always be recovered. A run against today’s corpus is not a reconstruction of yesterday’s evidence. Identify that limit accurately and avoid presenting a new execution as proof of the original event.

The part that is genuinely a judgement call

Hosted and self-hosted serving have different operating costs, skills and control requirements. Compare them against data boundaries, sustained utilisation, latency, availability and assurance needs. A hosted endpoint may suit some workloads, while local serving can be justified where its complete operating conditions are supported.

Location and operational-control requirements should be scoped to the workload. Evaluate an alternative serving arrangement before treating it as available. A gateway can help decouple some interfaces, but it does not make the entire hosting decision costlessly reversible.

What this does not tell you

This is a shape, not a product list. It says nothing about which gateway, which store or which serving stack — those depend on the estate, the skills on the ground and what the organisation already runs, and a reference architecture that names vendors is a procurement document wearing a diagram’s clothes.

It also does not claim that every system needs all five layers on day one. A single team with one use case and no regulatory exposure can reasonably start with three. What it should not do is start with three and never write down which two are missing, because that omission is the thing that becomes a rebuild.

The architecture owner should review the five concerns before accepting broader service use. Name the evidence, owner and limitation for each. That record lets the organisation judge the next use case against a defined capability rather than infer production readiness from a successful demonstration.

Filed under · Architecture · Inference · Architecture · Platform Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.