TL;DR
Firms selling services built on large language models cannot quote a stable price, because they cannot predict what the underlying compute will cost them. Goldman Sachs expects businesses to consume 24 times more tokens by 2030 than in 2026, reaching 120 quadrillion a month, even as the price per token falls. Volume is outrunning the discount.
The root problem is that identical inputs do not produce identical costs. Vary a prompt slightly and the answer changes; run the same prompt twice and it may still change; swap models and everything changes. Chain several agents together and the unpredictability compounds at every hop.
Simon Gooch of identity firm Saviynt put the commercial consequence plainly — committing a customer to a cost model spanning the next two or three years makes no sense when nobody knows what the inputs will cost. Will Venters of the London School of Economics frames it as a valuation problem rather than a budgeting one: output that is non-deterministic produces value that is non-deterministic too.
The failure mode is visible at large firms. Microsoft has reportedly restrained engineers’ use of certain third-party coding tools, and Uber is said to have consumed a year’s coding token allowance within months. That echoes Atlassian, which capped internal AI spending last week as bills climbed.
Smaller organisations are quietly exploiting the gap. Oliver King-Smith of smartR AI observes that they can stay beneath notice on flat-fee personal accounts, though he expects that to end once shareholder pressure for profit arrives and the platforms tighten up. For UK SMEs currently running on consumer-tier subscriptions, that is a repricing risk sitting on the books unrecognised.
Rob Steele, finance chief at UK accounting software firm iplicit, argues the discipline has to come from precision — you would not send a family member to do the weekly shop without telling them what to bring back. Costs scale badly once AI is embedded in a product used by thousands, and Venters notes that tokens get consumed by testing, security and guardrails as well as the core work. Adding agents takes a click, where adding staff takes a hiring conversation.
Bill Peterson of Sumo Logic, whose firm is previewing agentic security services, concedes nobody has solved it. The options are raising prices generally, charging on results, or bundling — and any of them can be upended when a model provider changes its own pricing. Customers, as he notes, cannot budget against something that moves every couple of months.
Looking Forward
OpenAI cut small-model prices last week, which helps buyers and does nothing for predictability. Outcome-based pricing is the obvious destination, but it requires vendors to carry the volatility themselves — and few have balance sheets for that.