OpenAI Built Its Own AI Chip. Here Is What That Means for Your Wallet and Your Wait Times.
OpenAI's custom silicon, called Jalapeño, promises faster responses and lower energy bills. The company plans to roll it out in small numbers before the end of 2025, with a bigger push in 2027.

Key points
- OpenAI's Jalapeño chip, built with semiconductor company Broadcom, delivers 1.5 to 1.9 times more AI work per unit of electricity than comparable Nvidia superchips in benchmark tests.
- The chip cuts end-to-end response time by 1.7 to 3.6 times across three large AI models tested.
- OpenAI will ship Jalapeño in small quantities by the end of 2025, scaling up through 2027.
- Nvidia remains part of OpenAI's hardware plan; Jalapeño does not replace it.
OpenAI has a new chip, and it is not from Nvidia.
The company published a blog post Tuesday describing Jalapeño, a custom processor built in partnership with Broadcom. The Verge AI first reported details from a briefing where OpenAI hardware vice president Richard Ho walked reporters through the numbers.
What is Jalapeño, and why should you care?
Jalapeño is an ASIC, short for application-specific integrated circuit, which simply means a chip designed to do one job extremely well rather than a broad range of tasks. That one job is AI inference: taking a finished AI model and running it fast enough to answer your question, summarise your document, or power an AI agent, meaning software that can carry out multi-step tasks on its own.
Every time you send a message to ChatGPT and wait for a reply, inference is happening. Faster, cheaper inference means shorter waits and, over time, lower costs for OpenAI to serve you.
How big is the performance gap?
OpenAI tested Jalapeño on InferenceX, a public benchmarking platform that scores how efficiently chips run AI models. The comparison was against Nvidia's GB200 and GB300 superchips, the current standard for this kind of work.
| Metric | Jalapeño vs. Nvidia GB200/GB300 |
|---|---|
| AI work per watt (GPT-OSS 120B) | 1.5x to 1.9x better |
| AI work per watt (DeepSeek R1) | 1.5x to 1.9x better |
| AI work per watt (Kimi K2.5 1T) | 1.5x to 1.9x better |
| End-to-end response time | 1.7x to 3.6x faster |
Ho said the chip breaks a common trade-off in AI hardware: systems usually force engineers to choose between speed and throughput, which is the volume of requests a chip can handle at once. Jalapeño, OpenAI claims, is strong on both.
What does this mean for people who use OpenAI's products?
Faster inference is the direct benefit. If OpenAI's claims hold at scale, ChatGPT responses could arrive more quickly, and AI agents built on the platform could feel more reactive.
Cheaper to run does not automatically mean cheaper for subscribers. But it does mean OpenAI can handle more users without proportionally higher electricity bills, which matters as AI demand keeps climbing.
Still, be cautious. These numbers come from OpenAI's own briefing, based on benchmark tests. Real-world performance across millions of varied requests is a different thing. The chips ship in "small volumes" by late 2025 and ramp up in 2027, so most users will not feel any change for a while.
OpenAI also made clear that Nvidia is not going anywhere. Ho described Nvidia as one of several "very good partners" in the company's overall computing strategy. Custom chips and bought chips will run side by side.
Honest takeaway: The chip itself is not something you buy or configure. But if the performance claims survive real-world deployment, the practical payoff is simpler: less waiting, more reliable access when demand spikes. Watch for independent benchmark results in 2026, when the rollout reaches meaningful scale.



