OpenAI unveils performance results for Jalapeno, its first custom inference chip, and says it improves both speed and power efficiency for AI model serving. In a Tuesday announcement and accompanying blog post, OpenAI says tests show higher throughput per watt and lower end-to-end latency than comparison systems, aiming to deliver “more intelligence from every watt” and faster responses.

Across reports, outlets describe the same core testing claims: Jalapeno is evaluated with OpenAI and third-party models (including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T) using a public benchmark approach from SemiAnalysis called InferenceX. OpenAI reports Jalapeno delivers roughly 1.5–1.9× more AI work per watt at peak throughput, and 1.7–3.6× lower end-to-end latency. Bloomberg and The Verge focus on the implication for deployment choices and user experience, while CNBC and Forbes emphasize the business and hardware-race angle of efficiency gains.

Several outlets also note boundaries to the comparison. OpenAI positions Jalapeno as inference-focused and not designed for training, and some reports add that Jalapeno was not tested against Nvidia’s newer Vera Rubin hardware. OpenAI says power use is rated at 700 watts with measured sustained power at or below 550 watts on tested workloads, and it plans to begin deploying Jalapeno in its infrastructure by the end of the year.