Just over one week after Nvidia agreed to backstop up to $105 billion in financing for its data centers, OpenAI arrived at Hot Chips on Tuesday with benchmarks claiming its first in-house chip beats Nvidia's GB300. Jalapeño, the inference ASIC OpenAI co-developed with Broadcom, delivered 1.5 times to 1.9 times more throughput per kilowatt and 1.7 times to 3.6 times lower end-to-end latency than Nvidia's GB200 and GB300 rack systems on SemiAnalysis's public InferenceX suite, with a 700W part going up against accelerators rated at 1,200W and 1,400W. OpenAI plans to begin deploying the chip in its own data centers later this year.
The tests covered three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5, with OpenAI reporting its widest leads at low-latency operating points, where it claims 8.6 times to 104.3 times more throughput per kilowatt at the GB300's fastest previous time-between-tokens settings.