On August 25, OpenAI unveiled its in-house inference chip Jalapeno at Hot Chips 2026. Co-developed with Broadcom and designed specifically for inference, the chip took only about 16 months from team formation to tape-out. Semiconductor research firm SemiAnalysis was invited to bring its benchmark suite to OpenAI's lab for testing and concluded that this first-generation chip can beat Nvidia Blackwell in almost all scenarios, achieving industry-leading levels in low latency, high throughput, and low concurrency.

Generally, first-generation in-house chips lack competitiveness, which is why the market is cautious about AI giants building their own chips. But SemiAnalysis's measurements show that Jalapeno has a thermal design power of only 700W, far below the 1,200W to 1,400W of Blackwell flagship cards. Its per-kilowatt inference throughput reaches 1.5x to 1.9x that of GB200 and GB300 rack systems, and end-to-end inference latency is 1.7x to 3.6x lower. The benchmarks cover multiple mainstream open-source models, including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5, and correctness evaluations are on par with Nvidia chips.

Interestingly, Nvidia's stock rose 2.19% on the day the benchmark results were released. On one hand, Nvidia also launched new products such as the Jetson Orin Nano 2 robotics computer. On the other hand, OpenAI explicitly stated that it will not abandon Nvidia, that training still relies heavily on it, and that the two companies also cooperate on financing guarantees. This signals that in-house chips will prioritize inference cost reduction, while the training ecosystem will remain Nvidia-led in the short term. Compute supply is moving toward stratification rather than replacement.
Of