OpenAI’s Jalapeño Chip Beats Nvidia Blackwell on Efficiency
2 min readOpenAI’s first piece of custom silicon is no longer a research curiosity. In benchmarks published this week, the company’s inference chip, code-named Jalapeño, delivered more AI work per watt than Nvidia’s Blackwell systems while drawing roughly half the power.
Why OpenAI built its own chip
Frontier AI labs have spent years renting or buying Nvidia accelerators, and Nvidia’s margins reflect that dependence. OpenAI began designing its own part with Broadcom to break the pattern, targeting inference specifically: the job of running a trained model for millions of users, which now consumes far more compute in aggregate than training does.
Jalapeño does not train models. It only serves them. That narrower scope is exactly what lets a custom design strip out silicon it does not need and spend the savings on efficiency.
What the benchmarks show
According to figures reported by CNBC and analyzed by SemiAnalysis, Jalapeño produced 1.5 to 1.9 times more output per watt at peak throughput across three tested models, with end-to-end latency 1.7 to 3.6 times lower than the best commercially available systems. The package draws 700 watts. Nvidia’s GB300 draws 1,400.
Analysts added an important caveat. Jalapeño uses newer HBM4 memory, so the Blackwell comparison is not like for like. Measured against Nvidia’s Vera Rubin platform, which also uses HBM4, Jalapeño still squeezes out more tokens per megawatt, though total cost of ownership per token lands roughly even between the two.
Why it matters
Power, not chip supply, is becoming the binding constraint on AI deployment. Data centers are being sited around available electricity, and a part that halves the wattage for the same output effectively doubles what an operator can serve from a fixed grid connection. That is a strategic result, not just a benchmark win.
It also puts pressure on Nvidia’s pricing power at the exact moment the company is reporting earnings. Google has run TPUs for years and Amazon has Trainium, but OpenAI is Nvidia’s most visible customer. A credible in-house alternative from that particular buyer changes the negotiating table. Watch for whether OpenAI scales Jalapeño beyond its own workloads, and whether Nvidia’s Rubin generation closes the efficiency gap.
Custom silicon has moved from hedge to leverage. The next question is how much of OpenAI’s inference fleet actually runs on it.
