OpenAI’s Jalapeño Chip Outpaces Rivals in AI Inference Performance
Iris Coleman
Aug 28, 2026 14:47
OpenAI’s Jalapeño custom AI chip achieves up to 1.9x better throughput per kilowatt and 3.6x lower latency than competing systems, redefining efficiency in AI inference.
OpenAI has unveiled benchmark results for its first custom AI inference chip, Jalapeño, showcasing industry-leading performance in efficiency and speed. According to OpenAI’s data, Jalapeño delivered between 1.5 to 1.9 times higher throughput per kilowatt and up to 3.6 times lower latency when compared to NVIDIA’s GB300-class systems. These gains highlight a significant leap in AI inference capabilities, particularly for large language models (LLMs) like GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T.
Jalapeño was first announced on June 24, 2026, as a collaboration with Broadcom, marking OpenAI’s move into custom silicon development. Unlike general-purpose GPUs, Jalapeño is purpose-built for running already-trained AI models, such as those powering ChatGPT. The chip is optimized specifically for inference workloads, minimizing power consumption and reducing response times. While NVIDIA’s GPUs dominate the broader AI hardware market due to their programmability and ecosystem, Jalapeño’s targeted design offers OpenAI a competitive advantage for its internal needs.
The benchmarks, released on August 25, highlight Jalapeño’s ability to handle interactive AI workloads with unprecedented efficiency. For example, on the largest tested model, Kimi K2.5, the chip achieved 1.5 times higher peak performance per watt and reduced end-to-end latency by 3.4 times compared to the competition. These metrics underscore its capability to process high-demand tasks, such as real-time chatbot interactions, with reduced energy costs — a crucial factor as AI adoption scales globally.
OpenAI’s engineering team credited AI itself for accelerating Jalapeño’s development. Leveraging internal AI tools, the team moved from design to tapeout in just nine months. Additionally, AI played a direct role in optimizing the chip’s circuits and programming, resulting in faster deployment and improved performance. Notably, AI-generated implementations for specific model blocks outperformed human-written code by 1.5 to 1.8 times, further streamlining development cycles.
Though Jalapeño is not available for external sale, its impact on OpenAI’s operations could be profound. Faster, more power-efficient inference allows the company to lower costs and serve more users, improving its operating leverage. With Gen 2 and Gen 3 chips already in development, OpenAI is doubling down on custom silicon as a strategic advantage.
In the broader market, Jalapeño’s performance positions OpenAI as a potential competitor to Google’s Tensor Processing Units (TPUs), though the two differ in scope. Google TPUs cater to both training and inference and are available commercially via Google Cloud, while Jalapeño is strictly an internal tool for inference. The comparison to NVIDIA is equally nuanced: while NVIDIA GPUs remain the go-to for versatility and ecosystem support, Jalapeño’s specialization offers superior efficiency for specific workloads.
Looking ahead, OpenAI plans to deploy Jalapeño at scale within its compute infrastructure by the end of the year. As the company continues to refine the platform and expand its capabilities, Jalapeño represents a key step in meeting the growing global demand for AI-powered applications while managing costs and environmental impact.
Image source: Shutterstock
