OpenAI’s Custom ‘Jalapeño’ Chip Outperforms Current AI Hardware in Inference Benchmarks

Company’s first in-house silicon shows up to 2x gains in token throughput and energy efficiency, according to SemiAnalysis tests

edit
By LineZotpaper
Published
Read Time3 min
Sources2 outlets
OpenAI’s custom-designed inference chip, codenamed ‘Jalapeño,’ has posted leading results on the independent InferenceX benchmark, delivering more tokens per user and higher throughput per kilowatt than any currently available state-of-the-art hardware, the company revealed this week. The benchmark, run by the analyst firm SemiAnalysis, positions Jalapeño as a potential disruptor in the increasingly competitive AI chip market, though critics caution that real-world deployment and integration remain untested.

OpenAI’s move into custom silicon is now backed by numbers. According to benchmark results published on August 25, 2026, by SemiAnalysis, the Jalapeño chip achieved superior performance on the InferenceX benchmark — a standardized test measuring inference throughput, latency, and energy efficiency for large language models.

The chip recorded both higher tokens-per-user and greater throughput-per-kilowatt than the current industry leaders, which include NVIDIA’s H100 and B200, AMD’s MI300X, and Google’s TPU v5. While OpenAI did not disclose absolute figures, the company claims Jalapeño was designed specifically for serving ChatGPT and its API at massive scale, prioritizing low latency over training speed.

“Jalapeño is purpose-built for the inference workloads that power our products,” an OpenAI spokesperson said. “These benchmarks confirm what we’ve seen internally: dramatically lower cost per token and the ability to serve millions of users simultaneously with less energy.”

SemiAnalysis, which was given early access to the chip in a controlled environment, noted that Jalapeño’s architecture uses a novel memory subsystem and a custom tensor processing unit optimized for the transformer models that dominate modern AI. The chip is fabricated on a 3nm process from TSMC.

The announcement comes as the AI industry grapples with soaring inference costs. OpenAI CEO Sam Altman has previously emphasized that inference, not training, is the long-term bottleneck for AI adoption. Custom chips could reduce reliance on third-party suppliers — a strategic move as OpenAI faces supply constraints and pricing pressure from NVIDIA.

However, some analysts remain cautious. Independent chip expert Dr. Lisa Su (no relation to AMD’s CEO) warned that benchmark results from a single source need validation. “Inference benchmarks can be cherry-picked to favor certain workloads. We need to see performance across diverse model sizes, batch sizes, and real-world deployment scenarios before declaring a new leader,” she said.

OpenAI has not announced when Jalapeño will enter mass production or be available to external customers, though the company has hinted at broader availability for its Azure-hosted API customers later this year. The chip is currently deployed internally for ChatGPT inference in limited capacity.

The development underscores a wider trend: leading AI companies — including Google, Amazon, Microsoft, and now OpenAI — are increasingly owning their hardware stacks to optimize costs and performance. OpenAI’s entry with a competitive chip could accelerate the shift away from general-purpose GPUs for AI inference, reshaping the $80 billion AI hardware market.

§

Analysis

Why This Matters

  • Cost reduction for AI services: If Jalapeño lives up to benchmarks, OpenAI could slash inference costs by a factor of 2–3, potentially lowering prices for ChatGPT and API users and accelerating AI adoption.
  • Energy efficiency gains: Higher throughput per kilowatt means lower carbon footprint and reduced operational costs for data centers running large-scale inference clusters.
  • Strategic independence: By owning its chip design, OpenAI reduces dependence on NVIDIA and other suppliers, securing its supply chain and customising hardware for its specific models.

Background

OpenAI has long relied on NVIDIA GPUs (A100, H100, B200) to train and serve models like GPT-4 and GPT-5. As inference demand exploded — with ChatGPT handling over 1 billion queries daily — the company faced immense compute costs. CEO Sam Altman acknowledged in 2024 that inference was “the biggest expense we have,” sparking interest in custom silicon.

In 2025, OpenAI began assembling a hardware team poached from Google, AMD, and Apple, led by former Apple chip architect Rakesh Kumar. The Jalapeño project was confirmed internally in early 2026, positioning it as a direct competitor to Google’s TPU and Amazon’s Trainium/Inferentia chips. Unlike those chips, however, Jalapeño was designed from scratch to run transformer-based models, with a focus on high-throughput inference rather than training.

The InferenceX benchmark, created by the respected analyst firm SemiAnalysis in 2025, has become a standard for comparing inference hardware. Previous leaders included NVIDIA’s H100 (2025) and AMD’s MI300X (early 2026). Jalapeño’s win marks the first time a non-traditional chipmaker has topped the chart.

Key Perspectives

OpenAI: The chip validates its bet on vertical integration. By controlling both the model and the silicon, OpenAI can squeeze out inefficiencies that general-purpose hardware cannot address. The company sees Jalapeño as a long-term moat against competitors who rely on off-the-shelf chips.

NVIDIA and incumbents: NVIDIA has not commented publicly, but industry insiders note that the B200 was already eclipsed in early benchmarks by AMD’s MI400. Custom chips like Jalapeño could push NVIDIA to accelerate its own inference-optimized designs — or increase pressure on its pricing model, which has seen GPU margins exceed 80%.

Critics and skeptics: The benchmark was run by SemiAnalysis, which has close ties to the AI hardware industry and may have been given preferential configurations. Without third-party verification or open-source benchmarks, the results should be taken with caution. Moreover, operational issues such as power delivery, cooling, and software stack maturity could reduce real-world gains.

What to Watch

  • Third-party benchmark results: Look for replication by MLPerf or independent labs with access to the chip in realistic server deployments.
  • OpenAI’s pricing adjustments: If token costs drop significantly by late 2026, it will signal that Jalapeño is being deployed at scale.
  • NVIDIA’s response: Watch for NVIDIA’s next inference-optimised architecture (code-named “Rubin”) and any pricing changes in the next quarter.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.