Give inference cost an owner and a unit of work
Inference spend depends on provider pricing, workload and architecture. Measure cost per accepted unit of work so model choice, context, retries and review effort can be assessed together.
Trace the interfaces between applications and records before changing a system.
Model spend that was small during a pilot can become material at production volume. The budget owner needs to understand which tasks generate the cost and what completed work the spending supports. A provider rate card gives only part of that explanation.
Assess pricing alongside workload and design. Model selection, request volume, context length and repeated calls all influence inference spend. A cheaper model may reduce cost or create additional review and rework. Compare options at an equivalent service and acceptance level rather than assuming either procurement or architecture is the sole cause.
Inference cost needs an accountable owner and a unit of work. Architecture choices and commercial terms should be evaluated together, with quality and human effort included. The purpose of measurement is to support a deliberate operating decision at the intended scale.
Where the money is actually decided
Four design choices are useful starting points for the cost review. Their relative importance depends on the workload and serving contract.
- Decision 01 How much context is sent Measure retrieval depth, chat history, prompts and tool schemas under the service’s pricing and cache rules.
- Decision 02 How many calls a task takes A single completion, a chain of five, or an agent loop with no fixed ceiling.
- Decision 03 Which model handles which request One model for everything, or a route that sends the easy majority somewhere cheaper.
- Decision 04 How the work is packed Compare batching, cache reuse and utilisation where supported by the serving system and contract.
Context length and call count can change cost without changing the model name. Measure their contribution on actual routes. An agent loop also needs enforced limits on steps, elapsed time and spend, with a defined escalation when it cannot finish within them.
Evaluate model routing on the organisation’s task and acceptance criteria. Smaller or cheaper models may be sufficient for some requests, but their suitability needs a local comparison. A market price comparison can identify candidates, while current provider terms establish the price and restrictions of the service being considered.
For owned or reserved compute, compare useful throughput with the full operating cost. Hardware, staffing, power and idle time can affect the result. Batching and cache reuse may improve utilisation where supported, including in some hosted services. Their value depends on workload shape, latency requirements and the contract.
The version of this that reaches a board
Translate infrastructure usage into the business unit: a completed case, ticket, claim or document. For work requiring acceptance, also report the share that meets the agreed standard. Monthly totals remain useful for budgeting, but they do not show whether unit economics improved.
Price the accepted decision
Swipe or scroll for the full diagram →
- Fixed assumptions · monthly GBP
- Fixed £500.00; model £0.04 per decision; loaded labour £30.00 per hour; rework share 0.05, cost £4.00 per reworked decision.
- Matched incumbent scenario
- £6,000.00 per month at the same volume; £0.75 per accepted decision at the same assumed acceptance share. Incumbent unit cost: £0.60.
- Accounting boundary
- Review is the initial human check. Rework is a separate correction after that check; its unit cost excludes the initial review labour. Both can occur on the same case.
- Formula and limit
- Monthly cost = fixed + volume × model unit cost + volume × review share × review minutes ÷ 60 × hourly labour + volume × rework share × rework unit cost. Divide by volume × acceptance share. This sensitivity model establishes neither causality, realised savings nor a staffing reduction.
Illustrative assumptions, not a price benchmark or savings forecast. Compare the incumbent at the same volume, acceptance and period.
Adjust monthly volume, review share, review minutes and acceptance share. All money is GBP per month.
Reviewed 2026-10-02
The calculator separates inference, review, rework and fixed costs using labelled assumptions. It compares the incumbent at the same volume, acceptance share and period. Replace the assumptions with measured inputs before using its result for an investment decision.
What to do before the next renewal
Set enforceable loop and spend limits, then compare model routes and context policies on the evaluation set. Assess any proposed saving with acceptance, latency, review and failure measures. There is no universal order for optimisation beyond preventing uncontrolled exposure and preserving the required service.
Also test whether the task needs a generative model. A statistical or rules-based baseline may meet the requirement with lower operating complexity. Its cost is not zero, and the comparison should include maintenance and error consequences alongside inference charges.
What this does not tell you
None of this is a forecast of what any organisation will pay. We publish no benchmark cost per case, because the number is entirely a function of the estate it applies to, and a figure quoted without that context becomes a target that somebody hits by degrading the system.
Cost measurement should help the owner choose an appropriate model and operating policy. More capable models may be justified for some tasks, while others may meet the requirement with a simpler route. Compare total cost and accepted outcomes before changing the allocation.
Before the next planning round, estimate the workload at the proposed volume and document the quality, review and capacity assumptions. Use the result to decide which routes can scale, what needs a different design and which commitments should wait for better evidence.