OpenAI’s custom Jalapeño inference chip outperforms Nvidia Blackwell

OpenAI is expanding beyond models and products into the chips and hardware that power them. Its first custom inference chip, Jalapeño, was built in just nine months through a partnership with Broadcom.

Early benchmark data released at the Hot Chips 2026 conference shows it outperforming Nvidia’s Blackwell in specific inference tasks. The company says Jalapeño is the first generation of its inference hardware, with generation 2 already in full development and generation 3 taking shape. If OpenAI eventually moves a part of its inference workload onto its own chips, it could reduce the cost of serving AI. As the team said, “Jalapeño gives us greater control over how our models run and over the economics of serving them.”

Jalapeño vs Blackwell

OpenAI and Broadcom report that during early internal testing Jalapeño delivers significantly higher performance-per-watt on LLM inference tasks compared to leading commercial hardware, including Nvidia’s flagship Blackwell architectures (like the GB200/GB300).

Given that data centers operate under strict electrical power constraints, higher performance per watt means OpenAI can process more AI tasks (tokens) for the same electricity bill. However, this does not mean Jalapeño replaces Blackwell. While Jalapeño is a specialized ASIC engineered strictly for running models, Blackwell remains the superior general-purpose platform for training them. The real economic shift is that OpenAI will no longer rely entirely on Nvidia GPUs for its massive, daily inference workloads.

At the same time, researchers at SemiAnalysis pointed out that comparing Jalapeño directly to Blackwell isn’t entirely fair because of hardware generation differences. Jalapeño relies on newer, ultra-fast HBM4 (High Bandwidth Memory 4) memory, whereas Nvidia’s Blackwell systems rely on older HBM3e. Analysts note that a truer competitor is Nvidia’s HBM4-equipped Vera Rubin platform, with perf-per-watt numbers are much closer.

Jalapeño’s performance

OpenAI tested Jalapeño with three very different models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Across all three models, Jalapeño achieved 1.5–1.9× more AI processing per watt at peak throughput than the systems used for comparison. It also reduced end-to-end latency by 1.7–3.6×, meaning the time to receive a response.

The following chart provided by OpenAI compares the energy efficiency of Jalapeño (green bars) against the existing best AI hardware accelerator (blue bars) across three AI models: GPT-OSS (535.28 tokens/s/user), DeepSeek R1 (169.41 tokens/s/user) and Kimi K2.5 (182.46 tokens/s/user). Jalapeño processes mixed tokens much faster while matching the top speeds of current accelerators.

Jalapeño energy efficiency compared to existing best accelerators (source: OpenAI)

While the previous chart focused on power efficiency (tokens per second per kilowatt), the next chart focuses on speed per user (low latency response generation). It compares the Jalapeño chip (green bars) to the leading hardware accelerator (blue bars) across three large language model sizes: GPT-OSS (120B parameters), DeepSeek R1 (670B parameters), and Kimi K2.5 (1T parameters).

Peak per-user decoding speed: Jalapeño compared to current best accelerators (source: OpenAI)

For applications where people are actively interacting with an AI, such as chatbots, coding assistants, or real-time agents, Jalapeño achieved 2.1–4.1× higher performance.

The next chart shows that Jalapeño consistently outperforms existing state-of-the-art accelerators by 1.5× to 1.9× in overall power efficiency.

Mixed-token efficiency across best operating points (source: OpenAI)

This metric, mixed tokens per second per kilowatt, indicates how many words or word-pieces (tokens) the chip can process for each unit of electricity used (kilowatt). Higher numbers are better because they represent greater efficiency. The chart shows that even when the current top chip is set up for its absolute best performance, the new Jalapeño chip still processes 1.5 to 1.9 times more data for every bit of electricity used across all model sizes.

These three charts demonstrate reduced operational costs and greater independence from GPU suppliers, such as Nvidia’s GB200/GB300 systems.

However, key contextual details should be added. Initial engineering samples were demonstrated in mid-2026. Limited deployment is expected by late 2026, with a broader data center rollout planned through 2027. OpenAI remains heavily reliant on Nvidia for model training and bulk compute.

Why you can’t rent Jalapeño

You cannot buy or rent Jalapeño because it is a proprietary, in-house chip designed exclusively for OpenAI’s internal data centers. It is not a commercial hardware product sold or rented to third parties. Similar to Google’s Tensor Processing Units (TPUs) or Amazon’s Trainium and Inferentia chips, Jalapeño was co-designed by OpenAI and Broadcom specifically to run OpenAI’s internal AI models and ChatGPT infrastructure.

How you benefit: Although you cannot buy or rent it, developers and end users benefit indirectly through faster response times. Jalapeño’s delivers much faster text generation speeds when using OpenAI services and models.

Our thoughts on Jalapeño

OpenAI’s new Jalapeño chip demonstrates that the company is expanding its focus to include custom hardware development. Because serving AI models to millions of daily users is costly, switching from standard Nvidia GPUs to custom ASICs could significantly reduce costs. While OpenAI will continue to use Nvidia for model training, developing its own chips provides greater control over costs, performance, and future technological advancements.

Jalapeño also reflects a wider industry moving toward custom AI hardware. Big tech companies are now building their own hardware to fit their exact needs and save energy. Even though Nvidia remains essential for model training, custom ASICs could play an important role in delivering efficient AI services at scale.

Read more:

Other popular posts