
OpenAI’s Jalapeño Inference Chip Is Staying In-House for Now
Published by AINave Editorial • Reviewed by Ramit
OpenAI’s Jalapeño inference chip is being deployed internally, not offered as an external hardware product. Hardware chief Richard Ho says the company expects to spend “a good long time” meeting its own growing compute demand, while leaving open the possibility of broader use later. The immediate story is less about a chip launch than about OpenAI building capacity for itself. Ho described internal deployment as the priority.
Running other models is a useful distinction
OpenAI designed Jalapeño around its own compute needs, with efficiency as a central goal. But the company says the chip is programmable and not hard-coded for OpenAI models. At Hot Chips, it demonstrated Jalapeño running GPT-OSS, DeepSeek R1 and Kimi K2.5. The demonstrations included both OpenAI’s model and models from other developers.
That flexibility matters because a custom accelerator can be closely matched to particular workloads. Jalapeño’s ability to run those additional models shows it is not limited to OpenAI’s own model family. It does not, by itself, show how easily the chip would support a broad range of customers or workloads. Ho also described codesign with OpenAI’s models as part of the approach, noting that model research IP can be difficult to share with outside silicon vendors. He said efficiency drove the design.
Published comparisons cover one test shape
OpenAI compared Jalapeño with Nvidia’s GB200 and GB300 using SemiAnalysis’ InferenceX benchmark. The reported tests used an 8k1k setup: 8,000 input tokens and 1,000 output tokens. Ho said OpenAI chose Blackwell because those were the best published results it could find. The comparison was limited to that fixed input and output length.
That makes the benchmark informative about a defined inference workload, not a general verdict on which accelerator is faster or more efficient across every deployment. Ho said OpenAI had also run internal tests with larger context windows and Nvidia’s Vera Rubin platform, but did not publish those results. He said Vera Rubin would be relevant by the time Jalapeño is deployed and characterized the internal comparison favorably. Those comments are OpenAI’s account, not public benchmark data. The Vera Rubin results remain unpublished.
Supply, not a customer launch, sets the near-term focus
Ho said OpenAI is in good shape on its own supply, but described meeting external demand as a separate challenge. He left the door open to wider use without offering a timeline or confirming an external rollout. OpenAI’s stated priority is to meet its own compute needs first.
For now, Jalapeño’s significance lies in what it could let OpenAI optimize internally, not in a new option for infrastructure buyers. Whether its workload-specific efficiency and ability to run other models can translate into a broader product remains open, especially while the company expects to keep the hardware busy itself.



















