Tech

OpenAI’s Jalapeño chip outperforms Nvidia Blackwell in inference benchmarks

New data from the Hot Chips conference reveals OpenAI’s custom silicon delivers superior throughput and efficiency, setting the stage for a multigenerational hardware strategy.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · View original source
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
Markets and Finance

OpenAI has presented its first batch of benchmark results for the Jalapeño chip, a custom inference processor developed in close collaboration with Broadcom. At the Hot Chips conference on Tuesday, the company shared detailed performance data indicating that Jalapeño achieves higher tokens per user and greater throughput per kilowatt than current state-of-the-art processors.

The benchmarks, conducted using Semianalysis’s InferenceX suite, positioned Jalapeño favourably against Nvidia’s Blackwell systems. Richard Ho, OpenAI’s head of hardware, described the findings as a “very, very significant performance advance” over the existing state of the art. He noted that the chip allows OpenAI to serve more AI work per unit of power while simultaneously returning responses more quickly, balancing high volume efficiency with low latency.

A key architectural feature of Jalapeño is its approach to minimising delays during the prefill and communication phases of processing, which OpenAI identifies as common bottlenecks. By keeping model state, including the KV cache used during response generation, local, the system reduces data movement and communication delays. This design allows the chip to activate the appropriate combination of compute, memory, and networking for each specific inference phase.

OpenAI intends for Jalapeño to become a multigenerational platform, coordinating the development of its AI products, models, chips, and memory. This full-stack approach enables the company to address specific friction points in the inference process that standard off-the-shelf hardware may not optimise for. Notably, OpenAI’s own models assisted in the development process of the chip, which was first announced in October last year.

Despite the strong benchmark results, investors should note that the comparison is against currently available Nvidia Blackwell systems. By the time Jalapeño reaches full deployment, the competitive landscape in silicon may have advanced significantly. Ho estimated that deployment would begin in “very small volumes” at the end of 2026, with more significant rollout expected in 2027.

The timing of the release places Jalapeño in a competitive environment where other major technology firms are also advancing their silicon roadmaps. For instance, Apple has recently announced M5 Ultra and M6 chips targeting on-device AI and distributed inference capabilities. As OpenAI moves from development to deployment, the efficiency gains demonstrated in these benchmarks will be critical for managing the rising costs of serving large-scale AI workloads.

Continue reading

More from Tech

Read next: Sam Altman rules out OpenAI IPO filing in 2026
Read next: Deep-tech startups dominate investor picks from Y Combinator’s latest Demo Day
Read next: Larry Ellison cancels planned US$7.5 billion Oracle share sale