Include power availability in the architecture review
Power availability can affect where AI capacity is offered, when it becomes available and what a workload costs. Include it in placement and capacity reviews, while keeping energy estimates separate from measured usage.
Trace the interfaces between applications and records before changing a system.
An AI architecture review should consider the infrastructure needed to serve the expected workload. Electricity supply and grid connections are part of that dependency, particularly when an organisation needs substantial capacity in a specific region. Their relevance extends beyond environmental reporting.
For a buyer of hosted inference, the practical questions concern availability, lead time and price. Residency obligations may restrict the regions that can serve the workload. Check whether the supplier can meet the required throughput and timing in those regions rather than assuming that capacity is available in every listed location.
Treat power availability as one input to architecture alongside performance, security, latency and cost. Its importance varies by workload and location. It does not replace the other constraints or imply that every AI deployment is limited by electricity supply.
What the published numbers actually say
The International Energy Agency’s work on energy demand from AI is the most useful public reference here, because it separates data centre demand overall from the AI-specific portion and states its uncertainty rather than hiding it. Its central projection has global data centre electricity consumption roughly doubling by 2030, with accelerated servers, mainly driven by AI, accounting for almost half of the projected net increase — and the report is candid that the spread between its scenarios is wide, which is itself the planning-relevant fact.
The IEA projections provide context for capacity planning, but do not determine a particular supplier’s price or availability. Ask providers for commitments relevant to the intended region and workload. Assess the consequence of unavailable capacity and distinguish those contractual facts from wider demand scenarios.
Where it lands in a design
- Decision 01 Placement Which regions can serve this workload, and whether the residency-constrained ones have capacity when you need it.
- Decision 02 Capacity commitment Reserved throughput versus on-demand. Reservation is now a hedge against availability, not only against price.
- Decision 03 Work per request How much reasoning, retrieval and re-ranking each request performs, and which work can be avoided.
- Decision 04 Batch versus interactive Compare scheduled work with interactive serving on turnaround, capacity and total cost.
- Decision 05 Model size Compare models against the task requirement, serving configuration and available efficiency evidence.
Batch processing is worth evaluating for work that does not require an immediate answer, such as scheduled classification or enrichment. It may offer different prices and capacity options, depending on the provider. Compare turnaround requirements, reliability and total cost before changing the serving path. Batching is an architectural option, not evidence of an energy saving.
The reporting problem underneath it
Estimating workload energy requires information about the model, hardware, utilisation and serving behaviour. Providers may not expose enough of it for a measured per-request figure. Published estimates should be read with their assumptions, including context length and reasoning behaviour. An average for an unspecified query is insufficient for a local comparison.
Tokens, request counts, model routes and execution regions can usually be recorded internally. These describe workload activity and help compare operating choices. Token counts alone do not measure energy and cannot establish a defensible emissions total without a validated conversion method and an appropriate electricity basis.
What this does not tell you
This article does not provide a measured energy figure for a deployed system. It uses published scenarios to identify planning questions. Any local estimate should state its method, assumptions, period and uncertainty, and should not be presented as a measurement when the necessary instrumentation is absent.
Efficiency per request and total consumption are different measures. Lower unit cost or resource use may support more requests, so a more efficient system can still consume more overall. Report both the workload volume and the unit estimate when assessing the effect. Additional capacity or flexibility can be valuable without establishing a climate benefit.
The architecture owner should review regional capacity at selection and renewal, particularly where residency constraints limit alternatives. Record the provider’s commitment, the demand assumption and the fallback. Revisit the decision when workload growth or supplier availability changes.