OpenAI’s First In-House Chip ‘Jalapeño’ Outperforms Nvidia GB300 in Efficiency Benchmarks

Co-developed with Broadcom, the 700W ASIC claims up to 1.9x throughput per kilowatt and 3.6x lower latency, deploying later this year

edit
By LineZotpaper
Published
Read Time3 min
OpenAI unveiled benchmarks for its first in-house inference ASIC, Jalapeño, at the Hot Chips conference on Tuesday, claiming it outperforms Nvidia’s flagship GB300 in throughput per kilowatt and latency by significant margins. The chip, co-developed with Broadcom in a nine-month design cycle, is slated for deployment in OpenAI’s own data centers later this year, marking a strategic shift toward vertical integration amid a tight HBM memory market.

OpenAI has released the first public benchmarks for Jalapeño, its custom inference accelerator, claiming it exceeds Nvidia’s GB300 in efficiency and latency across several large language models. The results, presented at the Hot Chips conference, demonstrate a 1.5x to 1.9x improvement in throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency compared to Nvidia’s GB200 and GB300 rack systems, according to tests run on SemiAnalysis’s public InferenceX suite.

The benchmarks covered three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s 1-trillion-parameter Kimi K2.5. OpenAI’s largest efficiency gains came at low-latency operating points, where it claimed up to 104.3 times more throughput per kilowatt than the GB300’s fastest previous time-between-tokens settings.

OpenAI normalized the results to each accelerator’s published package TDP (thermal design power), though the company noted that Jalapeño’s measured sustained power remained at or below 550W during testing—well below its 700W rating. An appendix using all-in utility power per accelerator narrowed the performance gap: 1.18kW for Jalapeño versus 2.55kW for the GB300. The gap also shrank when comparing against GB300 systems running multi-token prediction, a technique commonly used in production Nvidia deployments.

SemiAnalysis, which said it ran the benchmarks alongside OpenAI engineers in the company’s lab, described Jalapeño as “beating every Nvidia, AMD, and Google chip we have been able to test.”

Crucially, Jalapeño is an inference-only chip and does not handle training workloads, where Nvidia remains dominant. It was also not tested against Nvidia’s upcoming Vera Rubin platform, which OpenAI has agreed to deploy in a gigawatt-scale system starting in the second half of 2026.

Each Jalapeño package pairs its compute die with six stacks of HBM4 memory, totaling 216 GiB at 15.4 TB/s—roughly 50% more memory per watt than the GB300’s 288 GB of HBM3E at 1,400W. OpenAI stated that its architecture targets exposing aggregate HBM bandwidth rather than simply adding more memory, addressing a critical bottleneck in AI inference.

The chip’s efficient memory design comes at a time when HBM supply is extremely tight. Samsung, SK hynix, and Micron have sold out their HBM capacity through 2027, a shortage so severe that Nvidia is reportedly testing reduced memory configurations for its Rubin Ultra platform.

§

Analysis

Why This Matters

  • OpenAI’s move to in-house silicon reduces reliance on Nvidia and signals a broader industry trend toward vertical integration among AI labs.
  • The efficiency gains could lower operational costs for running large models, potentially making AI inference more affordable and accessible.
  • HBM memory shortages continue to roil the chip industry, affecting roadmap timelines and product configurations for both incumbents and new entrants.

Background

For years, Nvidia has dominated the AI accelerator market with its GPU-based solutions, creating a near-monopoly in both training and inference. OpenAI, one of Nvidia’s largest customers, has long relied on Nvidia hardware to train and deploy its models. However, growing costs and supply constraints—particularly for high-bandwidth memory (HBM)—prompted OpenAI to develop its own inference ASIC in partnership with Broadcom.

The design cycle was remarkably fast: just nine months from RTL to tapeout. Jalapeño was unveiled in June 2026, and now with published benchmarks, OpenAI is demonstrating its capabilities publicly for the first time. The company plans to deploy the chip in its own data centers before the end of the year.

The HBM market is currently in a state of severe shortage. Top-tier suppliers have sold out capacity through 2027, forcing even Nvidia to consider lower-memory SKUs for its next-generation chips. This supply crunch has become a strategic vulnerability for the entire AI industry.

Key Perspectives

OpenAI: The company positions Jalapeño as a major step toward optimizing AI inference for its specific workloads. By controlling the hardware, OpenAI can tune the architecture for latency and efficiency, reducing costs and improving user experience. The company emphasizes that Jalapeño targets inference only, and it continues to invest in Nvidia’s training hardware.

Nvidia: While not officially commenting on the benchmark comparison, Nvidia maintains that its GPUs (including the GB300) are general-purpose accelerators capable of both training and inference, whereas Jalapeño is a specialized ASIC. Nvidia may also argue that multi-token prediction and other production optimizations narrow the efficiency gap, and that upcoming Vera Rubin platform will reclaim performance leadership.

Industry Analysts and Skeptics: Some note that the benchmarks were run with OpenAI engineers present, raising questions about impartiality. The comparison at single-token prediction may not reflect real-world deployments where multi-token prediction is standard. Additionally, the chip’s long-term viability depends on continued access to HBM and Broadcom’s manufacturing capacity.

What to Watch

  • Deployment timelines: Will OpenAI meet its target of deploying Jalapeño in its own data centers by the end of 2026?
  • Nvidia’s response: How will Vera Rubin perform relative to Jalapeño, and will Nvidia adjust its memory configurations in response to the HBM shortage?
  • HBM supply dynamics: Any new capacity coming online from Samsung, SK hynix, or Micron could ease constraints and affect the competitive landscape.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.