On August 25, OpenAI publicly unveiled its first in-house inference chip, Jalapeno, at the Hot Chips 2026 conference. Semiconductor research firm SemiAnalysis, invited to OpenAI's labs for hands-on testing, concluded that the first-generation chip—with a 700W thermal design power—defeats NVIDIA's most powerful Blackwell series in nearly every scenario, achieving 1.5 to 1.9 times the per-kilowatt inference throughput of the GB300 rack system and delivering end-to-end inference latency up to 3.6 times lower.
First-generation self-developed chips are rarely expected to be competitive out of the gate, but Jalapeno went from team formation to tape-out in just roughly 16 months, completing tape-out in November 2025 with engineering samples already in hand. Benchmark testing covered multiple mainstream open-source large language models, including GPT-OSS 120B, DeepSeek R1 670B, and Moonshot Kimi K2.5. The chip achieved state-of-the-art results in both low-concurrency and high-throughput scenarios, while matching NVIDIA's accuracy on the GSM8k benchmark.

What deserves the most attention here isn't the chip's raw performance, but what it signals about the computing power supply landscape. OpenAI has made clear it will not abandon NVIDIA—training workloads will continue to rely heavily on NVIDIA hardware—but inference workloads are shifting toward energy-efficient in-house silicon. As leading model developers begin taking inference compute into their own hands, the computing power hardware market is moving from single-vendor dominance toward a diversified, multi-architecture landscape.
To be sure, the caveats should be stated clearly. The above results are based on single-token prediction; when compared against the GB300 with multi-token prediction enabled, the peak energy-efficiency advantage narrows to approximately 1.5 times. Moreover, Jalapeno remains an engineering sample, with volume production and gradual ramp not expected until 2027. Software stack maturity and large-scale deployment reliability remain unproven, meaning enterprises cannot simply adjust procurement plans based on these results alone.

For most enterprises, the more practical question is how to use computing resources effectively in an era of diversification. StarWar Cloud focuses on GPU computing power platforms and intelligent scheduling capabilities, helping enterprises manage training and inference workloads under unified governance with on-demand elastic allocation. This enables flexible integration of diverse computing power foundations—including domestic Chinese chips—providing agent-driven businesses with more headroom for computing power selection.
Looking deeper, Jalapeno's development process itself serves as a textbook example of AI-assisted chip design: the design phase leveraged Codex and GPT-Astra, with AI-generated kernels executing 1.5 to 1.8 times faster than those written by human experts. When AI begins participating in the design of computing hardware itself, both the iteration speed and supply elasticity of the entire computing power supply chain stand to be redefined.
From NVIDIA to OpenAI, competition in computing power hardware is evolving from single-player dominance to a multi-player dance. For application-side enterprises, computing power is no longer a fixed cost locked to a single vendor, but a resource network that can be orchestrated by scenario. Those who can integrate diverse computing power effectively will be best positioned to take the lead in the next wave of AI deployment.