OpenAI used its Hot Chips presentation to pull back the curtain on Jalapeño, the custom AI inference accelerator first unveiled in June. The chip, built with Broadcom, is a massive reticle-sized ASIC with 216 GB of HBM4 memory and up to 15.4 TB/s of bandwidth. It delivers up to 3.4 MXFP8 PFLOPS and 13.4 MXFP4 PFLOPS at a 700W power envelope, running at 1.70 GHz on silicon already in OpenAI's labs, with plans to push clocks to 1.80 GHz.
The company said a 128-accelerator Jalapeño domain provides more than 1 PB/s of aggregate HBM4 bandwidth, which it contrasts with Nvidia's GB300 NVL72 at 576 TB/s. For a one-trillion-parameter model with FP4 weights — roughly 0.5 TB — OpenAI argues the system could theoretically read the entire model more than 2,000 times per second. But the company acknowledged that real inference performance falls far short of theoretical bandwidth limits, which is why the design emphasizes what it calls a memory-sliced, NUMA-style spatial architecture rather than simply adding more HBM.
On paper, Jalapeño's raw specs look modest next to Nvidia's Blackwell Ultra accelerators (10 FP8 PFLOPS and up to 20 NVFP4 PFLOPS). OpenAI contends, however, that raw compute and memory bandwidth are not the differentiators. Instead, the company is prioritizing data movement and low-latency inference, claiming real-world wins over Nvidia's GB200 and GB300 in performance-per-watt and latency-sensitive workloads. The MXFP4 format, the company noted, may not be sufficient for training, positioning Jalapeño squarely as an inference engine.
The system scales to 128 accelerators in a local rack over Ethernet at 600 GB/s per processor, and up to 2,048 ASICs in a 16-rack pod configuration at 200 GB/s per processor. That full pod offers 27 exaFLOPS of MXFP4 compute, 432 TB of HBM4 memory, and 32 PB/s of aggregate memory bandwidth. Networking relies on Broadcom Tomahawk 6 Ethernet switches in a 'half-flattened' two-level Clos topology, with higher bandwidth for tensor-parallel traffic and lower bandwidth for expert-parallel communication. During the Q&A, OpenAI confirmed the scale-up network uses 200-Gb/s Ethernet links. Broadcom will manufacture the chip, while Celestica will build the surrounding hardware.
OpenAI's claims have not been independently verified, and the company has not published full benchmark methodology. The presentation did demonstrate working silicon, but real-world inference performance, particularly against Nvidia's next-generation platforms, remains to be validated in production deployments.