TL;DR
DeepSeek has formally launched V4 Pro at $1.32 per million tokens in and $3.96 out, roughly nine and fourteen times what the Flash tier charges. The company has also said it will lift API rates on both models and start charging differently at peak and off-peak hours. The firm that made cheap inference its reputation is now selling a premium product at premium rates.
The numbers
Artificial Analysis lists Flash at $0.14 and $0.28 for the same volumes, which is where the multiple comes from. On that firm’s Intelligence Index — a composite of nine measures covering long-context handling, coding, reasoning in the sciences, agentic work and the use of tools — the reasoning configuration of V4 Pro scores 53 against Flash’s 40.
That gap is the justification for the price. It is also a correction: Flash, launched last month, had embarrassingly outscored the April preview build of Pro in several independent tests, which is not the order these things are meant to arrive in.
Why UK buyers should care
Two divergent price moves landed on the same day. Google cut Gemini Flash to half its predecessor’s rate; DeepSeek moved its flagship sharply upward and flagged further rises to come. Anyone treating token costs as a reliably falling line has just had that assumption tested from both directions in twenty-four hours.
The peak and off-peak structure deserves attention too. It signals capacity constraint rather than a pricing experiment, and it hands cost control back to whoever can schedule batch work overnight. For UK firms running document processing, summarisation or agentic pipelines on a Chinese model to keep costs down, both changes belong in the next budget review — and the exercise is worth repeating across every vendor in the stack, not just this one.
The company behind it
DeepSeek is spending to stay in the race. It raised roughly $7.4bn in June, its first outside money, and by July was reportedly seeking more at about a $74bn valuation. It intends at least to double headcount, including data-centre and agent teams, and has been recruiting chip designers to build its own silicon and lean less on Nvidia and Huawei. Meanwhile ByteDance and Alibaba, along with MiniMax, Zhipu AI and Moonshot AI, have eroded the lead R1 won it in early 2025.
Looking forward
Higher prices are how that investment gets paid for. The open question is whether customers who came for cheap tokens stay for expensive ones.