On August 25, OpenAI presented the first official test results of its specialized inference chip, Jalapeño, and they look like a serious challenge to Nvidia's hegemony in the AI accelerator segment. According to my data, in the company's benchmarks, the new processor demonstrates 1.5 to 1.9 times greater computational efficiency per watt and a 1.7–3.6 times reduction in latency compared to flagship systems based on Nvidia GB200 and GB300.
For the most demanding interactive scenarios, the gap becomes even more impressive: Jalapeño's performance surpasses competitors by 2.1–4.1 times. OpenAI has already announced plans to deploy the chip in its own infrastructure by the end of 2026, which could radically change the economics of its cloud services.
Detailed breakdown of the tests: three models, three wins
For objectivity, OpenAI used the public InferenceX benchmark from SemiAnalysis, which evaluates the full request processing cycle, including throughput, power consumption, and latency. The results are impressive:
GPT-OSS 120B vs GB200: Jalapeño's peak performance was 85,448 mixed TPS/kW versus 44,960 for Nvidia (a 1.9 times advantage). Full latency was 1.03 seconds versus 1.8 seconds.
DeepSeek R1 670B vs GB300: here, the efficiency gap reached 19,641 mixed TPS/kW versus 11,781 (1.7 times), and latency was 3.6 times lower—1.65 seconds versus 5.99 seconds.
Kimi K2.5 1T vs GB300: OpenAI's proprietary chip delivered 18,195 mixed TPS/kW versus 11,862, achieving latency of 1.56 seconds versus 5.31 seconds. The advantage is 1.5 and 3.4 times, respectively.
An important nuance: the calculations used TDP figures declared by the manufacturers. For Jalapeño, this is 700 W, while the GB200 has 1200 W and the GB300 has 1400 W. At the same time, OpenAI notes that the chip's actual sustained power consumption in tests did not exceed 550 W, making the results even more compelling.
AI as an engineer: a nine-month development cycle
The methodology behind creating Jalapeño deserves special attention. OpenAI actively used its own AI models at all stages, from searching for architectural solutions to verifying and optimizing arithmetic blocks. This made it possible to shorten the path from concept to production-ready design to nine months—an unprecedentedly short timeframe for the semiconductor industry.
Moreover, using the Codex tool based on GPT-Astra, engineers managed to achieve high chip performance on three open-weight models—which were not originally part of the production plan—in less than two months. For individual components—attention and mixture-of-experts in GPT-OSS—AI-generated implementations proved to be 1.5–1.8 times faster than manually written ones.
Strategic context: more than just a chip
OpenAI's Chief Financial Officer, Sarah Friar, positions Jalapeño as a key element of a vertically integrated infrastructure covering data centers, memory, networks, and software. The proprietary accelerator gives the company unprecedented control over inference costs and allows flexible workload distribution based on price-to-performance ratios.
Notably, OpenAI is not abandoning partners—the portfolio includes Microsoft, Nvidia, AWS, AMD, Broadcom, and others. The economic logic here is simple: reducing the cost per unit of computation directly improves the ratio of revenue to infrastructure expenses. Friar also references Jevons' paradox, suggesting that increased efficiency will lead not to reduced compute consumption, but to explosive growth in the number of economically viable AI tasks.
Jalapeño is only the first generation. The second is already at an advanced stage of development, and the third is in the design phase. The first chips will appear in OpenAI's infrastructure before the end of the year.
My analytical conclusion: Jalapeño's success is not just a technological victory, but a signal of a paradigm shift in the industry. If OpenAI manages to scale production and confirm these metrics in real-world operation, we will witness the first serious challenge to Nvidia's dominance in the AI inference segment in recent years. The only question is whether OpenAI can maintain this advantage under conditions of mass production and growing demand.