OpenAI (Unlisted:OPAI) said Tuesday its custom-built Jalapeño inference chip, developed with Broadcom Inc (NASDAQ:AVGO, XETRA:1YD), delivered better throughput per watt and lower response latency than Nvidia Corp (NASDAQ:NVDA, XETRA:NVD)'s GB300 in internal testing.
The chip is built specifically for inference rather than model training and runs at approximately 700 watts. OpenAI plans to begin deploying it for its models later this year.
The performance advantage widened on larger workloads, including Moonshot's Kimi model, and Jalapeño has also performed well internally on unreleased OpenAI models, the company said. It was tested against Nvidia's GB300, not the newer Vera Rubin generation.
A second-generation chip is already approaching tape-out, with work on a third generation underway.
In testing, DeepSeek R1 exceeded 700 tokens per second per user at concurrency of one, while Kimi K2.5 and GPT-OSS reached roughly 1,400 tokens per second per user in selected tests, using single-token prediction without speculative decoding.
The chip is built on TSMC's N3P process with HBM4 memory offering 15.4 terabytes per second of bandwidth, which SemiAnalysis said is currently the highest among shipping or near-shipping accelerators. A next-generation version, B0, is already in fabrication and expected to improve performance per watt by roughly 25%.
Up to 2,048 Jalapeño chips can be connected across a 16-rack scale-up domain, with OpenAI's next deployment target at around 100 megawatts. Production is expected to ramp through 2027, with most output planned for the fourth quarter.