Wake up and the AI world has exploded. OpenAI's first self-developed inference chip, Jalapeno, officially released its measured data - the first-generation product directly surpassed Nvidia's current top-of-the-line Blackwell product family.
In SemiAnalysis's InferenceX benchmark, Jalapeno delivered a report card that shocked the entire semiconductor industry: AI workload per watt reached 1.5-1.9 times that of Nvidia's GB300, end-to-end latency dropped by 1.7-3.6 times, and interactive workload performance improved by 2.1-4.1 times. In high-concurrency scenarios with the DeepSeek R1 670B model, the per-watt throughput gap at the same decoding speed even reached the hundredfold level.
This chip went from architecture design to tape-out in only nine months. The hardware lead, Richard Ho, was once a core engineer on Google's TPU; combined with OpenAI's own large models deeply participating in chip design optimization, the industry's usual R&D cycle of two to three years was compressed by two-thirds.
Many people think Nvidia is finished, but things are far from that simple. Jalapeno is a pure inference-only ASIC - it does no training and is not sold externally, serving only OpenAI's internal use. Its core goal is to reduce the company's own inference costs and escape dependence on a single supplier. In other words, this is OpenAI's "private computing reserve" built for itself, not a general-purpose product to sell to the whole industry.
The moat of the CUDA ecosystem remains bottomless. Millions of developers worldwide, countless frameworks and toolchains are bound to Nvidia's ecosystem - this cannot be instantly overturned by one chip's high benchmark scores. What can truly shake Nvidia's position is a good-and-cheap alternative ecosystem, not merely a fast chip.
What is truly striking about this chip is the signal it releases: AI model companies building their own silicon has gone from rumor to mass-production-level capability. When the people who understand models best personally design the chips best suited to those models, the efficiency of this full-stack software-hardware collaboration is hard for traditional chipmakers to match.
The second-generation chip has entered late-stage development, and the third generation has started concept design. Within the next two to three years, OpenAI's inference costs will see a cliff-like drop, and GPT's response speed and concurrency will rise to a new level.
Nvidia remains the king, but challengers have risen beneath the throne. The second half of the computing-power war has only just begun.




