TL;DR
Britain is arguing hard about whether to build AI data centres and barely at all about whether the demand forecast underneath them holds. A contrarian case is forming in the trade press: as the industry’s centre of gravity moves from training models to running them, a large share of that work may never touch a hyperscale GPU cluster at all.
Where the argument comes from
The reasoning has two halves. First, training is a one-off cost concentrated in a few labs, while inference is the recurring workload that scales with adoption — and inference is far less demanding. Second, inference is migrating outward, onto laptops, phones and small local machines, with only the queries that exceed local capacity routed to central hardware.
Several trends reinforce it. Self-hosting an open-weight model, whether on kit you own or rented capacity, can substantially undercut frontier-model pricing. Specialised smaller models often outperform a general-purpose one on a narrow task, in the way an oncologist need not know astronomy. And silicon designed purely for inference is arriving: AMD acquired the Toronto startup Taalas in August 2026, whose approach bakes a model permanently into the chip and which claims speed improvements of a different order to a general-purpose GPU. Vendors making these claims are selling the alternative, so the numbers deserve scepticism — but the direction of travel does not depend on any single figure.
Why this matters for the British argument
Read this against the rest of today’s coverage and the awkwardness is obvious. The Greens want construction paused; the STUC wants an industrial strategy with enforceable job conditions; Labour rejects a pause on the grounds that jobs and sovereign capacity are at stake. Every one of those positions assumes the demand curve is real and steep.
If a meaningful share of inference ends up on devices, some of that capacity gets built for a load that never arrives — which changes what a community is being asked to accept in exchange, and how much leverage it holds. Our earlier reporting found grid connection requests tripling around Newport and Cardiff. Those queues are being sized against a forecast, not a measurement.
Looking forward
Sensible planning does not need this thesis to be right, only to be possible. Routing between models by cost and sensitivity is becoming standard practice for larger buyers, and every prompt handled locally is one that never reaches a rack in Fife or Buckinghamshire. For UK firms, the near-term implication is procurement rather than politics: paying frontier prices for work a smaller model handles is now the most common avoidable AI cost.