OpenAI’s Jalapeño Inference Chip Is Staying In-House for Now
tomshardware.com

OpenAI’s Jalapeño Inference Chip Is Staying In-House for Now

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DROpenAI is deploying its Jalapeño inference chip internally and expects its own growing compute needs to keep the hardware occupied for a long time. The chip has run models beyond OpenAI’s, but its public performance comparisons cover a specific test workload.

OpenAI’s Jalapeño inference chip is being deployed internally, not offered as an external hardware product. Hardware chief Richard Ho says the company expects to spend “a good long time” meeting its own growing compute demand, while leaving open the possibility of broader use later. The immediate story is less about a chip launch than about OpenAI building capacity for itself. Ho described internal deployment as the priority.

Running other models is a useful distinction

OpenAI designed Jalapeño around its own compute needs, with efficiency as a central goal. But the company says the chip is programmable and not hard-coded for OpenAI models. At Hot Chips, it demonstrated Jalapeño running GPT-OSS, DeepSeek R1 and Kimi K2.5. The demonstrations included both OpenAI’s model and models from other developers.

That flexibility matters because a custom accelerator can be closely matched to particular workloads. Jalapeño’s ability to run those additional models shows it is not limited to OpenAI’s own model family. It does not, by itself, show how easily the chip would support a broad range of customers or workloads. Ho also described codesign with OpenAI’s models as part of the approach, noting that model research IP can be difficult to share with outside silicon vendors. He said efficiency drove the design.

Published comparisons cover one test shape

OpenAI compared Jalapeño with Nvidia’s GB200 and GB300 using SemiAnalysis’ InferenceX benchmark. The reported tests used an 8k1k setup: 8,000 input tokens and 1,000 output tokens. Ho said OpenAI chose Blackwell because those were the best published results it could find. The comparison was limited to that fixed input and output length.

That makes the benchmark informative about a defined inference workload, not a general verdict on which accelerator is faster or more efficient across every deployment. Ho said OpenAI had also run internal tests with larger context windows and Nvidia’s Vera Rubin platform, but did not publish those results. He said Vera Rubin would be relevant by the time Jalapeño is deployed and characterized the internal comparison favorably. Those comments are OpenAI’s account, not public benchmark data. The Vera Rubin results remain unpublished.

Supply, not a customer launch, sets the near-term focus

Ho said OpenAI is in good shape on its own supply, but described meeting external demand as a separate challenge. He left the door open to wider use without offering a timeline or confirming an external rollout. OpenAI’s stated priority is to meet its own compute needs first.

For now, Jalapeño’s significance lies in what it could let OpenAI optimize internally, not in a new option for infrastructure buyers. Whether its workload-specific efficiency and ability to run other models can translate into a broader product remains open, especially while the company expects to keep the hardware busy itself.

FAQs

It is a custom inference ASIC designed around OpenAI’s compute needs, with efficiency as a central goal. OpenAI describes it as programmable rather than hard-coded for its own models. Ho outlined those design priorities.

Sources

Latest Tech News