TL;DR
OpenAI has released the first performance numbers for Jalapeno, the inference silicon it designed in-house, and claims it delivers between 1.5 and 1.9 times the work per watt of the Nvidia systems it was measured against, with latency between 1.7 and 3.6 times lower. Every figure comes from OpenAI itself and has not been independently verified.
What the company is claiming
The tests ran on InferenceX, a publicly documented benchmark published by SemiAnalysis, across three open models: GPT-OSS 120B, DeepSeek R1 at 670B and Kimi K2.5 at a trillion parameters. On the largest of those, the reported gain is about 1.5x at peak on the per-watt measure, with round-trip latency cut by a factor of 3.4. For highly interactive work — the agent-style pattern where delays compound step by step — it puts the gain between 2.1 and 4.1 times.
The power framing does a lot of work here. OpenAI normalised results against each accelerator’s published rating, which puts its own 700W part against Nvidia comparators rated at 1,200W and 1,400W, and notes Jalapeno drew 550W or less in sustained use. Performance per watt rather than per chip is a defensible metric for anyone paying an electricity bill, but it is also the metric that flatters the lower-rated part.
Two development details are worth noting separately from the benchmarks. OpenAI says AI tooling took the design from concept to tapeout inside nine months, and that model-written kernels outran human-expert implementations by 1.5 to 1.8 times on selected attention and mixture-of-experts blocks.
Why UK buyers should care
Resultsense reported on Monday that Nvidia server prices are set to rise by more than 15% as memory costs bite, and that OpenAI had cut frontier model prices by a fifth for three months. Those two stories pull in opposite directions, and this one explains how both can be true: a provider that owns its inference hardware can absorb a component squeeze its competitors have to pass on.
For UK organisations budgeting AI spend, that is the practical read. Falling API prices are not necessarily a discount funded by generosity or by loss-leading — they can be a structural cost advantage, which makes them more likely to persist.
Looking forward
Deployment starts on OpenAI’s own estate before this year closes, with a second generation already in development. The company says it will keep buying Nvidia regardless. Independent benchmarking, when it arrives, is what turns these claims into procurement evidence.