OpenAI details Jalapeño AI accelerator at Hot Chips, claims inference gains over Nvidia Blackwell

Custom ASIC co-developed with Broadcom reached tape-out in nine months; 2,048-chip system scales to 27 exaFLOPS

edit
By LineZotpaper
Published
Read Time3 min
At the Hot Chips conference, OpenAI disclosed new architectural details and performance figures for its Jalapeño inference processor, claiming the custom ASIC co-developed with Broadcom beats Nvidia's GB200 and GB300 in low-latency inference and performance-per-watt. The chip, which reached tape-out in just nine months, uses a NUMA-style spatial architecture and can scale to 2,048 accelerators delivering a combined 27 exaFLOPS of MXFP4 performance.

OpenAI used its Hot Chips presentation to pull back the curtain on Jalapeño, the custom AI inference accelerator first unveiled in June. The chip, built with Broadcom, is a massive reticle-sized ASIC with 216 GB of HBM4 memory and up to 15.4 TB/s of bandwidth. It delivers up to 3.4 MXFP8 PFLOPS and 13.4 MXFP4 PFLOPS at a 700W power envelope, running at 1.70 GHz on silicon already in OpenAI's labs, with plans to push clocks to 1.80 GHz.

The company said a 128-accelerator Jalapeño domain provides more than 1 PB/s of aggregate HBM4 bandwidth, which it contrasts with Nvidia's GB300 NVL72 at 576 TB/s. For a one-trillion-parameter model with FP4 weights — roughly 0.5 TB — OpenAI argues the system could theoretically read the entire model more than 2,000 times per second. But the company acknowledged that real inference performance falls far short of theoretical bandwidth limits, which is why the design emphasizes what it calls a memory-sliced, NUMA-style spatial architecture rather than simply adding more HBM.

On paper, Jalapeño's raw specs look modest next to Nvidia's Blackwell Ultra accelerators (10 FP8 PFLOPS and up to 20 NVFP4 PFLOPS). OpenAI contends, however, that raw compute and memory bandwidth are not the differentiators. Instead, the company is prioritizing data movement and low-latency inference, claiming real-world wins over Nvidia's GB200 and GB300 in performance-per-watt and latency-sensitive workloads. The MXFP4 format, the company noted, may not be sufficient for training, positioning Jalapeño squarely as an inference engine.

The system scales to 128 accelerators in a local rack over Ethernet at 600 GB/s per processor, and up to 2,048 ASICs in a 16-rack pod configuration at 200 GB/s per processor. That full pod offers 27 exaFLOPS of MXFP4 compute, 432 TB of HBM4 memory, and 32 PB/s of aggregate memory bandwidth. Networking relies on Broadcom Tomahawk 6 Ethernet switches in a 'half-flattened' two-level Clos topology, with higher bandwidth for tensor-parallel traffic and lower bandwidth for expert-parallel communication. During the Q&A, OpenAI confirmed the scale-up network uses 200-Gb/s Ethernet links. Broadcom will manufacture the chip, while Celestica will build the surrounding hardware.

OpenAI's claims have not been independently verified, and the company has not published full benchmark methodology. The presentation did demonstrate working silicon, but real-world inference performance, particularly against Nvidia's next-generation platforms, remains to be validated in production deployments.

§

Analysis

Why This Matters

  • OpenAI is moving from being a dominant AI model developer into custom silicon, a shift that could reshape the economics of AI inference and reduce its reliance on Nvidia.
  • The nine-month tape-out cycle — unusually fast for a reticle-sized ASIC — suggests that AI-assisted chip design and aggressive co-development with Broadcom are compressing hardware development timelines.
  • If Jalapeño's performance-per-watt claims hold up in production, it could put pressure on Nvidia's data-center GPU pricing and roadmap.

Background

OpenAI unveiled Jalapeño in June 2026, marking its first custom inference processor, developed in partnership with Broadcom. The chip was notable not just for its specs but for its development speed: reaching tape-out in nine months is exceptionally fast for a chip of this complexity. The architecture is built on a NUMA-style spatial design, which OpenAI says is optimized for moving model weights efficiently rather than maximizing raw peak FLOPS. The company has positioned Jalapeño as an inference-only part, with training still handled by GPUs. The broader context is an industry-wide push by hyperscalers and AI labs — including Google, Amazon, and Meta — to design custom accelerators to cut costs and improve efficiency for AI workloads.

Key Perspectives

OpenAI: Argues that low-latency inference and bandwidth efficiency matter more than brute-force compute, and that Jalapeño's Ethernet-based scale-out design offers flexibility and cost advantages over proprietary NVLink fabrics. The company claims superior performance-per-watt versus Nvidia's GB200 and GB300. Nvidia: Has not publicly responded to OpenAI's claims. Nvidia's position is that its mature CUDA ecosystem, full-stack software, and raw peak performance (Blackwell Ultra delivers roughly 10 FP8 PFLOPS per part) remain the industry standard for AI workloads, including inference. Critics and skeptics: Note that OpenAI's benchmarks are self-reported and lack independent verification. Questions remain about whether MXFP4-only peak performance limits training use cases, whether the 200-Gb/s Ethernet scale-up fabric will bottleneck tensor-parallel workloads at full pod scale, and whether the chip's physics — with a 700W power rating and reticle-sized die — will translate to real cost savings after packaging, memory, and system integration costs are factored in.

What to Watch

  • Independent benchmarks or production deployment reports that verify OpenAI's low-latency and performance-per-watt claims against Nvidia GB200/GB300 systems.
  • The transition from 1.70 GHz to 1.80 GHz clocks — whether the chip can hit higher clocks in mass production without thermal or reliability issues.
  • OpenAI's future roadmap: whether Jalapeño is followed by a training-capable ASIC, a next-generation inference part, or broader availability of the design for external customers.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.