Start a conversation Contact
← Research

A smaller model on your own hardware is a governance decision

Self-hosting an open-weight model is usually argued as a cost saving and bought as a sovereignty control. Both framings hide the thing that actually changes, which is who becomes responsible for behaviour that used to be somebody else's problem.

The proposal arrives with a spreadsheet. At current volume the hosted API costs this much, a pair of accelerators costs that much, the crossover is somewhere in month nine, and after that the organisation is saving money and its data never leaves the building. It is a good spreadsheet. Every number in it is defensible.

It is also answering a question the organisation has not asked, which is whether it wants to become the operator of a model rather than the customer of one. Those are different businesses with different obligations, and the second one is not obviously worse — for a growing number of European enterprises it is becoming the right answer — but it should be chosen deliberately rather than arrived at through a total cost of ownership calculation.

The claim: self-hosting moves a set of responsibilities across an organisational boundary, and the cost case does not price them.

What actually moves

Who holds what, on each side of the decision Fig. 01
Hosted endpoint Your own weights, your own hardware
The provider decides when a model changes, and tells you You decide. Nothing changes until you change it — including a fix you needed
Safety behaviour comes with the product, tuned by someone else Refusal behaviour, filtering and abuse handling are yours to build and evidence
Capacity is elastic and priced per token Capacity is what you bought. Utilisation is the number that decides the cost
Data leaves your boundary under a contract Data does not leave. The boundary is now something you have to prove
Vulnerabilities in the serving stack are patched for you The serving stack is software you run, on a patch cycle you own
Documentation for the model is whatever the provider publishes Documentation is what you can establish about weights you did not train

The second row is the one that surprises people. A hosted frontier model arrives with a large amount of work already done on what it will and will not produce. An open-weight model arrives with less of it, in a form the operator can modify, which is the point. If the system is customer-facing, that work has to exist somewhere, and after the decision it exists in your engineering plan and in your evidence pack.

The last row is the one that matters for anyone with regulatory exposure. A deployer of a hosted model can point at the provider’s documentation. An organisation that takes open weights, adapts them and puts them in front of a consequential decision has changed its role — and the obligations that attach to providing a system are not the obligations that attach to using one.

When it is the right answer anyway

Often. This is not an argument against it. It is an argument for reaching it through the right question.

The question that decides where inference should run Fig. 02

What is the binding constraint on this workload?

  • Data cannot cross a boundary — residency, contract, classification Self-host, and design the evidence that the boundary holds. The strongest reason, and the one that does not depend on volume.
  • Volume is high, stable and predictable Self-host if you can keep utilisation up. Otherwise you are buying idle hardware. The economics are a utilisation bet, not a hardware purchase.
  • The task is narrow and well specified A small model, hosted or not, probably beats a frontier model on cost and latency. Classification, extraction, routing and structured drafting rarely need the largest model available.
  • The work is open-ended and volume is spiky Stay hosted. The elasticity is the product. Research, long-form analysis, anything with a demand curve nobody can forecast.

The third row is the one most organisations should act on before the first. Model size and hosting are separate decisions that get bundled, and unbundling them is where the easy savings are. A large share of enterprise traffic is narrow work sent to a general model out of habit, and moving it to a smaller model — on anyone’s infrastructure — costs less and returns faster than a hardware programme.

The European version of this question

Residency has stopped being the whole of the question in Europe. The distinction being drawn in procurement now is between where data sits and who controls the stack it is processed on — jurisdiction over the operator, not only the postcode of the rack. The European Commission’s work on a common way to assess cloud and AI sovereignty is an attempt to make that assessable rather than rhetorical, and it is worth reading the framing directly on the Commission’s digital strategy pages before a supplier explains it to you.

For an organisation with genuine residency obligations, that shift changes the shortlist. An EU region operated by a non-EU provider answers one question. It does not answer all of them, and the ones it does not answer are the ones a regulator or a customer is most likely to ask.

What to establish before committing

What this does not tell you

We do not resell platforms and we do not take vendor commission, so this piece has no preference between the two answers. What it has is a strong preference for the decision being made against a stated constraint rather than a spreadsheet, because the spreadsheet is right about the hardware and silent about everything else in the table above.

It also does not claim that open-weight models are less capable than hosted ones for enterprise work. For a large proportion of the tasks organisations actually run, the published gap has narrowed to the point where it is not the deciding factor. The deciding factors are operational, and they are the ones a capability comparison will never surface.

The reader who acts differently is the one holding the crossover chart. Before approving it, write the binding constraint at the top of the page. If it is cost alone, the cheaper answer is almost always a smaller model on somebody else’s infrastructure — and that can be tested next week rather than next quarter.

Filed under · Architecture · Inference · Sovereignty · Architecture Inference Institute · 09 Jul 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.