OpenAI says its Jalapeño chip delivers 1.5‑1.9× more AI work per watt and 1.7‑3.6× lower latency than Nvidia chips across §4‑OSS, DeepSeek R1 and Kimi K2.5 1T
benchmarks chips deepseek inference nvidia openai
| Source: Techmeme | Original article
OpenAI reports its Jalapeño chip delivers 1.5‑1.9× more AI work per watt and 1.7‑3.6× lower latency than Nvidia chips on inference benchmarks across GPT‑OSS, DeepSeek R1 and Kimi K2.5 1T.
OpenAI has released the first performance figures for Jalapeño, its inaugural custom inference ASIC. In tests that spanned three distinct language‑model families – GPT‑OSS 120 billion‑parameter, DeepSeek R1 and Kimi K2.5 1T – the chip delivered between 1.5 × and 1.9 × more AI work per watt of power than Nvidia’s GB200/GB300 “Blackwell” superchips. The same benchmark suite, InferenceX, showed end‑to‑end latency reductions of 1.7 × to 3.6 ×, meaning responses are noticeably faster at comparable throughput.
The figures matter because inference efficiency has become a decisive factor in the economics of large‑scale AI services. More work per kilowatt lowers operating costs and reduces the carbon footprint of data‑center deployments, while lower latency directly improves user experience in products such as ChatGPT. By posting a clear advantage over Nvidia’s flagship offering, OpenAI signals that it can compete on the hardware front, potentially reshaping the balance of power in a market long dominated by Nvidia’s CUDA ecosystem.
What to watch next includes OpenAI’s rollout strategy for Jalapeño‑powered servers and whether the company will price the hardware competitively for external partners. Industry observers will also be looking for Nvidia’s technical response—whether a new generation of GPUs or software optimisations can close the gap. Finally, further independent benchmark releases will be crucial to validate OpenAI’s claims and to gauge how quickly the efficiency edge translates into real‑world cost savings for AI providers.
Sources
Back to AIPULSEN