OpenAI Ships Its First In-House Inference Chip: Jalapeño
OpenAI unveiled Jalapeño — its proprietary inference chip — which tops the InferenceX benchmark in tokens per user and efficiency per watt.
OpenAI unveiled Jalapeño — a proprietary chip designed specifically for running (inferencing) AI models, not training them. It is the company's first silicon solution: until now, OpenAI was entirely dependent on third-party GPUs, primarily Nvidia's.
Benchmark Results
According to an independent InferenceX test by analytics firm SemiAnalysis, Jalapeño outperforms all current market solutions on two metrics simultaneously: tokens per user (response speed) and tokens per watt (energy efficiency). OpenAI VP of Hardware Richard Ho called it "the best of both worlds" — low latency and high throughput at the same time.
What This Changes for the Industry
Building a proprietary chip is a strategic move: OpenAI reduces its dependence on Nvidia and gains control over token costs. If Jalapeño genuinely delivers better performance per watt than existing GPUs, it creates direct downward pressure on industry pricing. After Google TPU and Amazon Trainium, another major player has entered the custom AI silicon space.
For API users and developers, Jalapeño means potential token price reductions from OpenAI — the chip can serve more requests on the same power draw. Competition in AI hardware is no longer Nvidia's monopoly: the next two to three years in this segment look set to be eventful.
Source: openai.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles