Nvidia Is Pushing Hard to Own Every Chip Inside AI Data Centers
The company best known for graphics chips now wants to supply the brains that run AI agents too. Here is what its new Vera Rubin system actually does, and why it matters for anyone who uses AI tools at work.

Key points
- Nvidia's new Vera Rubin chip system claims to process ten times as many AI tasks per watt as its previous Grace Blackwell system.
- OpenAI already has one Vera Rubin rack running in its facilities, Nvidia confirmed at a technical briefing in June 2025.
- Nvidia is selling the Vera CPU, the processing brain inside the new system, as a standalone product and told Chinese customers it could ship as early as August 2025.
- AMD revealed its competing Helios AI chip rack on Sunday, the same week Nvidia staged its Vera Rubin briefing, in a clear race for multi-year chip contracts with companies like Meta, Amazon and OpenAI.
Nvidia built its fortune selling GPUs, the specialised chips that do the heavy number-crunching AI needs. Now it wants to sell the other kind of chip inside AI data centers too.
At a briefing for journalists at its Santa Clara headquarters last week, first reported in detail by Wired AI, Nvidia executives laid out the case for Vera Rubin, its next-generation chip system. The headline claim: Vera Rubin processes ten times as many tokens per watt as its predecessor. A token is a small chunk of text or data, and tokens per watt is essentially a measure of how much AI work a system can do for a given amount of electricity.
That matters to companies running AI at scale. Electricity is one of the biggest costs in a data center, and efficiency gains translate directly into lower bills.
Why is Nvidia selling CPUs now?
Because AI is getting more complicated, and that requires a different kind of chip. GPUs crunch numbers fast. CPUs, the general-purpose processors that have run computers for decades, are better at coordinating tasks, managing data flows and running software logic. As companies build AI agents, software that can carry out multi-step tasks on its own like booking a meeting or analysing a spreadsheet without a human clicking through each step, they need more CPU firepower alongside their GPUs.
Vera Rubin pairs one CPU with every two GPUs. A full Vera Rubin NVL72 rack, a stack of chips packed into a single liquid-cooled unit, holds 36 Vera CPUs and 72 Rubin GPUs.
Nvidia also claims the new system is far easier to install than earlier products. It describes Vera Rubin as "cable-free compute" and says setup time can drop from a couple of hours per rack to a few minutes. The whole system runs on liquid cooling, which uses less energy than blowing air over hot chips.
The company does benchmark the Vera CPU as faster than rival chips from AMD and Intel. Worth noting: the comparison tests apparently used slightly older generations of those competitors' products, which is the kind of fine print that always deserves attention when a company grades its own homework.
Early customers named by Nvidia include Microsoft, OpenAI and Oracle. Vera Rubin is expected to ship broadly in the second half of 2025.
For anyone whose job touches AI tools, the practical upshot is simpler than the chip specs suggest. Faster, more efficient hardware means the AI services you already use, from writing assistants to customer-service bots, get cheaper and quicker to run. That tends to make them more widely deployed, not less. Whether that is a threat or a gift depends almost entirely on how you use them.
Takeaway: If your organisation is evaluating AI tools this year, ask vendors which hardware generation their product runs on. Efficiency improvements at the chip level often show up as lower per-use costs within 12 to 18 months.



