OpenAI Built Its Own AI Chip. Here Is What That Means for Your Wallet and Your Wait Times.

OpenAI's custom silicon, called Jalapeño, promises faster responses and lower energy bills. Small volumes ship before the end of 2025, with a bigger push in 2027.

AI2Day NewsdeskAI-assistedPublished Updated Editor: Lee Brown3 min read
Illustration: A modern open-plan office desk seen from above at a slight angle
Illustration made with AI. Not a photograph of the events described.
Share

Key points

  • OpenAI's Jalapeño chip, built with Broadcom, delivers 1.5 to 1.9 times more AI work per unit of electricity than comparable Nvidia superchips in benchmark tests.
  • The chip cuts end-to-end response time by 1.7 to 3.6 times across three large AI models tested.
  • OpenAI will ship Jalapeño in small quantities by the end of 2025, scaling up through 2027.
  • Nvidia remains part of OpenAI's hardware plan; Jalapeño does not replace it.

OpenAI has a new chip, and it didn't come from Nvidia.

The company published a blog post Tuesday describing Jalapeño, a custom processor built with Broadcom. OpenAI hardware vice president Richard Ho walked reporters through the numbers at a briefing, with The Verge AI first to report the details.

What is Jalapeño, and why should you care?

Jalapeño is an ASIC, short for application-specific integrated circuit: a chip built to do one job extremely well rather than a broad range of tasks. That job is AI inference, the process of running a finished AI model fast enough to answer a query or power an AI agent, meaning software that carries out multi-step tasks on its own.

Every message you send to ChatGPT triggers inference. Faster, cheaper inference means shorter waits and, over time, lower costs for OpenAI to serve you. We've tracked this corner of the hardware race since August, including Nvidia's own latency work and a software-only bet from French startup Kog: OpenAI's move is the first time a major model provider has published benchmark data for its own silicon.

How big is the performance gap?

OpenAI tested Jalapeño on InferenceX, a public benchmarking platform that scores how efficiently chips run AI models. The comparison was against Nvidia's GB200 and GB300 superchips.

Metric Jalapeño vs. Nvidia GB200/GB300
AI work per watt (GPT-OSS 120B) 1.5x to 1.9x better
AI work per watt (DeepSeek R1) 1.5x to 1.9x better
AI work per watt (Kimi K2.5 1T) 1.5x to 1.9x better
End-to-end response time 1.7x to 3.6x faster

Ho said Jalapeño breaks a common trade-off: AI hardware usually forces engineers to choose between speed and throughput, the volume of requests handled at once. OpenAI claims its chip is strong on both.

What does this mean for people who use OpenAI's products?

Faster inference is the direct benefit. If the claims hold at scale, ChatGPT responses arrive sooner and AI agents feel more reactive.

Cheaper to run doesn't automatically mean cheaper for subscribers. It does mean OpenAI can handle more users without proportionally higher electricity bills, which matters as demand climbs.

These numbers come from OpenAI's own briefing, based on controlled benchmarks. Real-world performance across millions of varied requests is a different thing. Chips ship in small volumes by late 2025 and ramp in 2027, so most users won't feel any change for a while.

OpenAI also made clear Nvidia isn't going anywhere. Ho described Nvidia as one of several "very good partners" in the company's overall computing strategy.

Honest takeaway: The chip's real test isn't a benchmark deck; it's 2026, when the rollout reaches meaningful scale and independent results start coming in. Until then, the headline number to hold onto is that 3.6x latency improvement: if even half of it survives contact with production traffic, the waiting gets noticeably better.

© 2026 AI2Day