
OpenAI’s first custom AI chip has delivered higher performance per watt than Nvidia’s Blackwell systems in benchmark tests, giving the company a more efficient option for running large language models (LLMs).
The chip, called Jalapeño, handled between 1.5 and 1.9 times more AI work per kilowatt than Nvidia’s GB200 and GB300 systems across three models tested on the public InferenceX benchmark. It also recorded between 1.7 and 3.6 times lower end-to-end latency.
The results give OpenAI its first measured evidence that designing hardware around the workloads its models need can produce significant efficiency gains.
Jalapeño Takes on Nvidia’s Blackwell Systems
OpenAI developed Jalapeño with Broadcom as an application-specific chip designed for AI inference, which is the process of running trained models to generate responses. Unlike general-purpose accelerators, Jalapeño was built around the memory movement, networking, and computing patterns involved in serving LLMs.
On the benchmark results that were published using InferenceX, OpenAI compared Jalapeño with Nvidia’s GB200 for GPT-OSS 120B and with the GB300 for DeepSeek R1 670B and Kimi K2.5 1T.
On GPT-OSS 120B, Jalapeño produced about 85,448 mixed tokens per second per kilowatt, compared with 44,960 for the GB200. This gave the OpenAI chip about 1.9 times the peak throughput per kilowatt and 1.7 times lower end-to-end latency.
On DeepSeek R1 670B, Jalapeño reached 19,641 mixed tokens per second per kilowatt compared with 11,781 for the GB300. Its end-to-end latency was 1.65 seconds, compared with 5.99 seconds for Nvidia’s system.
And for Kimi K2.5 1T, Jalapeño delivered 18,195 mixed tokens per second per kilowatt against 11,862 for the GB300, with end-to-end latency lower at 1.56 seconds compared with 5.31 seconds.
The Chip Uses Less Power
Jalapeño has a 700-watt package power rating, compared with 1,200 watts for the GB200 and 1,400 watts for the GB300 used in the comparisons. OpenAI said its measured sustained power stayed at or below 550 watts on the workloads tested.
This efficiency matters as AI companies continue adding data center capacity to handle growing demand for inference. Also, more work from each unit of power can allow operators to serve more requests without increasing energy use at the same rate.
SemiAnalysis, which created InferenceX, also examined the results in person at OpenAI’s facilities. However, it noted that OpenAI supplied the performance figures and that it did not run the complete InferenceX suite or its preferred AgentX benchmark. This means the results are significant, but they do not establish that Jalapeño is faster or more efficient across every AI workload.
OpenAI Still Needs Nvidia
Jalapeño is designed for inference, so the results do not mean OpenAI has replaced Nvidia across its AI infrastructure. OpenAI continues to use Nvidia chips alongside hardware from other suppliers, while developing future generations of its own accelerators.
The comparison is also limited to Nvidia’s Blackwell generation, as Nvidia’s newer Vera Rubin platform was not part of these tests.
For now, Jalapeño shows what OpenAI can achieve when its models, software, and hardware are designed together. The bigger test will come as the chip moves from benchmark results into wider deployment and faces the next generation of AI hardware.
