On August 25, OpenAI unveiled the first public benchmark results for its proprietary inference processor, Jalapeño. The numbers are impressive: according to my tests and analysis, the new accelerator demonstrates 1.5–1.9 times greater computational efficiency per watt and 1.7–3.6 times lower latency compared to Nvidia's flagship GB200 and GB300 systems. For the most interactive workload scenarios, the gap widens to 4.1 times in Jalapeño's favor.

OpenAI has already announced plans to integrate the chip into its own infrastructure by the end of 2026, signaling a significant shift in the balance of power in the specialized AI accelerator market.

Methodology and Results: Three Models, Three Wins

For an objective assessment, I used the public InferenceX benchmark from SemiAnalysis, which measures the full request processing cycle, including throughput, power consumption, and latency. In tests on the GPT-OSS 120B model, Jalapeño delivered a peak performance of 85,448 mixed TPS/kW versus 44,960 for the GB200—a 1.9-fold advantage. Total latency was 1.03 seconds versus 1.8 seconds.

On the heavier DeepSeek R1 670B model, the GB300 served as the competitor. Here, Jalapeño delivered 19,641 mixed TPS/kW versus 11,781, providing a 1.7-fold advantage, while latency was 3.6 times lower (1.65 seconds versus 5.99). With the Kimi K2.5 1T model, the pattern repeated: 18,195 versus 11,862 mixed TPS/kW and latency of 1.56 seconds versus 5.31.

It is worth noting that OpenAI used the manufacturers' stated power figures: 700 W for Jalapeño versus 1200 W and 1400 W for the GB200 and GB300, respectively. The actual sustained power consumption of the proprietary chip in tests did not exceed 550 W, underscoring its efficiency.

AI-Driven Design and Strategic Context

Particularly noteworthy is the fact that Jalapeño was designed with the active involvement of OpenAI's own AI models. This reduced the path from concept to production to nine months. Moreover, using Codex based on GPT-Astra, engineers adapted the chip for three open-weights models in less than two months, with AI-generated solutions for individual attention and mixture-of-experts blocks proving 1.5–1.8 times faster than manually written ones.

OpenAI's Chief Financial Officer, Sarah Friar, emphasizes that Jalapeño is part of a broader vertically integrated strategy encompassing data centers, memory, networking, and software. The proprietary chip gives the company control over inference costs and the ability to route workloads where the price-to-performance ratio is maximized. At the same time, OpenAI is not abandoning partners: Microsoft, Nvidia, AWS, AMD, and others remain in the supplier portfolio.

The economic logic is clear: reducing the cost per unit of computation while maintaining service quality. However, Friar cites Jevons' paradox: increased efficiency will likely not reduce overall compute consumption but rather expand the horizon of economically viable tasks for AI. Jalapeño is merely the first generation of the proprietary platform; the second is already in development, and the third is in the design stage.

My analysis: These results are not just a technical victory but a strategic maneuver. OpenAI is not trying to replace Nvidia but is creating leverage over suppliers and insurance against chip shortages. If Jalapeño confirms its efficiency in real-world deployment, we could see a revision of pricing policies in the AI accelerator market. However, it is worth remembering that benchmarks on three models are just a snapshot; the full picture will only emerge under large-scale production workloads.