Request a scoping call Contact
← Research

A smaller model on your own hardware is a governance decision

Self-hosting can offer control over inference, but it also transfers operational responsibilities. Assess utilisation, serving security, evaluation and supplier evidence alongside the cost and residency case.

Architecture / Conceptual study
Make the connections explicit.
  1. Applications
  2. Interfaces
  3. Records

Trace the interfaces between applications and records before changing a system.

A self-hosting proposal may compare API charges with accelerator capacity and forecast a cost crossover. That comparison is useful, but the investment decision also needs to account for the responsibilities the organisation will assume as the operator of a model-serving system.

Model selection and hosting are separate choices. A smaller model might suit the task whether it runs on an external service or internal infrastructure. Self-hosting becomes appropriate when its control, performance or economic advantages justify the operating burden in the organisation’s setting.

Assess the transfer of operational responsibility alongside the hosting economics. The decision needs an owner for updates, serving security, capacity, evaluation and incident response. Include those activities in the comparison rather than assuming that a hardware purchase replaces the full service.

What actually moves

Who holds what, on each side of the decision Fig. 01
Hosted endpoint Your own weights, your own hardware
The provider decides when a model changes, and tells you You decide. Nothing changes until you change it — including a fix you needed
Safety behaviour comes with the product, tuned by someone else Refusal behaviour, filtering and abuse handling are yours to build and evidence
Capacity is elastic and priced per token Capacity is what you bought. Utilisation is the number that decides the cost
Data leaves your boundary under a contract Inference can remain inside the chosen boundary. Verify telemetry, dependencies and support access.
Vulnerabilities in the serving stack are patched for you The serving stack is software you run, on a patch cycle you own
Documentation for the model is whatever the provider publishes Documentation is what you can establish about weights you did not train

Safety behaviour depends on the selected model, its tuning and the serving product. Open weights do not establish either an absence of safeguards or their adequacy for the intended use. Evaluate the actual refusal, filtering and abuse-handling behaviour, and decide which controls the operator must maintain.

Regulatory roles require a separate assessment. Self-hosting or adapting weights does not automatically determine a legal role. Intended purpose, branding and the nature of a modification can matter. Article 25 of the EU AI Act identifies circumstances in which another party becomes the provider of a high-risk system. Establish the relevant facts with counsel before allocating obligations.

When it is the right answer anyway

Self-hosting can be appropriate where a documented boundary, predictable workload or need for operational control supports it. Test those conditions explicitly.

The question that decides where inference should run Fig. 02

What is the binding constraint on this workload?

  • Data cannot cross a boundary — residency, contract, classification Self-host, and design the evidence that the boundary holds. The strongest reason, and the one that does not depend on volume.
  • Volume is high, stable and predictable Self-host if you can keep utilisation up. Otherwise you are buying idle hardware. The economics are a utilisation bet, not a hardware purchase.
  • The task is narrow and well specified Compare smaller hosted and self-hosted candidates on quality, latency and total operating cost. Classification, extraction, routing and structured drafting rarely need the largest model available.
  • The work is open-ended and volume is spiky Compare hosted elasticity with the capacity and operating costs of self-hosting. Research, long-form analysis, anything with a demand curve nobody can forecast.

Test model size independently of hosting. For a narrow classification, extraction or routing task, compare a smaller candidate with the current approach on the same cases. Include quality, latency and total operating cost. A favourable result can reduce the capacity requirement before an infrastructure commitment is made.

The European version of this question

Residency has stopped being the whole of the question in Europe. The distinction being drawn in procurement now is between where data sits and who controls the stack it is processed on — jurisdiction over the operator, not only the postcode of the rack. The European Commission’s work on a common way to assess cloud and AI sovereignty is an attempt to make that assessable rather than rhetorical, and it is worth reading the framing directly on the Commission’s digital strategy pages before a supplier explains it to you.

For an organisation with genuine residency obligations, that shift changes the shortlist. An EU region operated by a non-EU provider answers one question. It does not answer all of them, and the ones it does not answer are the ones a regulator or a customer is most likely to ask.

What to establish before committing

What this does not tell you

We do not resell platforms and we do not take vendor commission, so this piece has no preference between the two answers. What it has is a strong preference for the decision being made against a stated constraint rather than a spreadsheet, because the spreadsheet is right about the hardware and silent about everything else in the table above.

Capability comparisons alone cannot decide this question. The relevant evidence is whether a candidate meets the workload’s quality and operating requirements under the proposed controls. Do not assume that an open-weight or hosted model will be adequate because it belongs to either category.

Before approving capacity, record the binding constraint and the evidence that supports it. If cost drives the proposal, compare smaller hosted and self-hosted candidates with matched quality and utilisation assumptions. If control drives it, require evidence that the proposed boundary covers telemetry, support access, dependencies and recovery as well as inference traffic.

Filed under · Architecture · Inference · Sovereignty · Architecture Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.