Jalapeño's inference gain does not yet end OpenAI's chip dependency
OpenAI's first Jalapeño measurements show higher throughput per kilowatt and lower latency than Blackwell on selected models. Small 2026 volumes, no external sales and continued reliance on other suppliers for training limit what those numbers establish today.
Artificial Intelligence··Morning
The benchmark compares inference
OpenAI says Jalapeño, developed with Broadcom, delivered more work per kilowatt and more tokens per user than an NVIDIA Blackwell system on SemiAnalysis's InferenceX benchmark. TechCrunch reports that the test used GPT-OSS 120B and that the figures have not been independently reproduced. Axios says the company also measured DeepSeek R1, the 1 trillion parameter Kimi K2.5 and OpenAI's open model. This is a comparison for selected inference workloads, not a result about training performance.[1], [2]
Deployment begins at small volume
Richard Ho said only a very small number of Jalapeño systems will enter service at the end of 2026, with broader capacity following in 2027. TechCrunch reports that the design keeps the key-value cache local to reduce prefill and communication bottlenecks. Axios says a second generation is already in development and plans for a third are taking shape. The first measurements therefore do not yet show how much of OpenAI's total inference traffic will move to its own chip this year.[1], [2]
Training remains the dependency boundary
OpenAI does not plan to sell Jalapeño to outside customers; Ho said the company's own demand leaves no room for that. Because the chip is not designed for training, OpenAI remains dependent on NVIDIA and other suppliers for that work. How far rival systems advance before broad deployment begins is unknown. Jalapeño's strategic effect will therefore be determined not by one benchmark ratio but by the share of OpenAI's inference workload it carries over time.[1], [2]