OpenAI publishes first measured results for JalapeΓ±o, its first custom inference chip: 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency than comparison systems
In a company engineering publication dated 25 Aug 2026, OpenAI released the first measured performance results for JalapeΓ±o, described as "OpenAI's first custom inference chip". Tested on InferenceX, "a public benchmark from SemiAnalysis", across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI reports JalapeΓ±o "delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems", and 2.1 to 4.1 times higher performance on highly interactive workloads. OpenAI states it plans "to begin deploying JalapeΓ±o within OpenAI's compute infrastructure by the end of the year" and that it is "the first generation of a multigenerational roadmap" with Gen 2 "deep in development". The chip was co-developed with Broadcom per OpenAI's earlier unveil publication. Material because it is working first-party silicon with measured third-party-benchmark results from the largest buyer of merchant AI accelerators, landing the day before Nvidia's quarterly print.