Tuesday, August 25, 2026

OpenAI’s Jalapeno chip is built for fast inference at scale, benchmarks show

Good news! OpenAI is also getting into the chip business.

Why did OpenAI opt to write Jalapeño with the Spanish special character instead of a clean and simple Jalapeno? Grrr! This special character is not on my keyboard! I hope, OpenAI will correct this mistake! Keep it simple stupid (KISS)!

"At the Hot Chips conference on Tuesday [8/25/2026], OpenAI shared a more detailed look at Jalapeno, including the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors. ...

First announced last October, Jalapeño was developed by OpenAI in close collaboration with Broadcom, with OpenAI’s own models assisting in the development process. The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory all developed in concert.

Because of that full-stack approach, OpenAI was able to address specific phases in the inference process that often cause friction during inference processing. In particular, Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says often act as bottlenecks.

“We designed Jalapeno to minimize data movement and communication delays,” ..."

"... OpenAI models also accelerated Jalapeno’s development. Earlier generations helped the team design and bring up the chip, while our latest models are accelerating how we optimize and program it. Jalapeno’s performance extends across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T, showing that the architecture works across models developed both inside and outside OpenAI.
Across all three, Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance. ..."

OpenAI’s Jalapeno chip is built for fast inference at scale, benchmarks show | TechCrunch







No comments: