OpenAI released performance benchmarks for its first custom AI inference chip, Jalapeño. The company asserts the hardware surpasses all currently available commercial systems.

The chip delivers 1.5 to 1.9 times more AI work per watt than competitors. It achieves 1.7 to 3.6 times lower end-to-end latency. Performance gains reached even higher levels for highly interactive workloads.

OpenAI tested the chip using GPT-OSS, DeepSeek R1, and Kimi K2.5 1T models. Broadcom co-developed the hardware to support OpenAI's shift toward a full-stack strategy. This move aims to reduce the company's reliance on third-party chipmakers like Nvidia.

OpenAI plans to deploy the chip within its infrastructure by the end of the year.