TL;DR: DeepSeek released V4-Flash on Friday, and research firm Artificial Analysis rates it the cheapest well-known model to run by a wide margin. It costs roughly 3 cents per benchmark test. Moonshot’s Kimi K3 comes in at 86 cents, OpenAI’s GPT-5.6 Sol at $1.86, and Claude Fable 5 from Anthropic at $3.15. On headline pricing, a million input tokens costs $0.14 and a million output tokens $0.28.
Cost per completed task is the more honest comparison, and it is the one Artificial Analysis used. Token pricing alone can mislead: a cheap model that needs several more steps to arrive at an answer can cost more in practice than a pricier one that gets there directly.
Cheap, but not equivalent
On capability the picture is less flattering. V4-Flash scored 50 out of 100 on the firm’s Intelligence Index, which aggregates nine benchmarks across coding, reasoning and workplace-style tasks. That matches Google’s Gemini 3.6 Flash and sits a point below both Meta’s Muse Spark 1.1 and Z.AI’s GLM-5.2. Kimi K3 reached 57, while Claude Opus 5, Claude Fable 5 and GPT-5.6 all came in nine or more points clear.
So the choice is not cheaper-for-the-same. It is roughly two-thirds of the frontier score at around one per cent of the running cost.
What that changes for UK buyers
For a good deal of production work — classification, extraction, summarisation, routing — a score of 50 is sufficient, and a hundredfold cost difference stops being a procurement detail and becomes an architectural one. The sensible pattern is tiering: route the bulk of volume to cheap models and reserve frontier capability for the calls that genuinely need it. That requires knowing which of your workloads are which, and most organisations have not measured it.
Looking forward
DeepSeek is reportedly preparing for a possible flotation and is trying to regain ground after Chinese rivals including Moonshot, MiniMax, Z.AI, ByteDance and Alibaba overtook it in attention. A more capable V4-Pro is planned with no announced date. Alibaba separately unveiled Qwen3.8-Max on Monday, its largest model yet. The competitive dynamic among Chinese labs is pushing inference costs down faster than Western pricing has moved, and UK firms building on per-token economics should be re-running their assumptions rather than treating last year’s costs as fixed.