NVIDIA's New Vera Rubin Chip Does 30 Times More AI Work Per Watt Than Its Predecessor
A new benchmark puts NVIDIA's Vera Rubin NVL72 hardware far ahead on energy efficiency for the kind of complex, multi-step AI tasks that companies are rapidly building into their products.

Key points
- NVIDIA's Vera Rubin NVL72 system delivers up to 30 times higher AI throughput per megawatt than the previous GB300 NVL72, according to NVIDIA's own measurements.
- Agentic AI workloads, where software carries out long chains of tasks automatically, consume 15 times more computing resources than a simple chatbot conversation, per OpenRouter data.
- Token costs, the charge companies pay for every word an AI model processes or produces, drop by up to 35 times on Vera Rubin NVL72 compared with GB300 NVL72.
- Results are pending independent review by research firm SemiAnalysis and do not yet include performance from the Vera CPU component.
- Vera Rubin is in full production and shipping to customers.
What is an AI agent, and why does it eat so much power?
An AI agent is software that works through a long chain of steps on its own, rather than just answering a single question. Think of it less like a search engine and more like a junior analyst who can open spreadsheets, call other tools, and write up findings without being asked twice at each stage.
That independence is expensive. According to data from OpenRouter, a service that routes traffic to AI models, agentic tasks use 15 times more tokens than a plain chat message. Tokens are the small chunks of text, roughly three-quarters of a word each, that AI models read and write. More tokens means more chips, more electricity, and bigger bills.
When an agent researches a company for an investment report, for example, it queries databases, reads news filings, spins up a second agent to model valuations, and stitches everything together. Every step feeds its output into the next, so the pile of text the model must hold in memory grows continuously. Some sessions run to hundreds of thousands of tokens.
What did NVIDIA actually announce?
The numbers are large. NVIDIA says its new Vera Rubin NVL72 system delivers up to 30 times more agentic work per megawatt of electricity than the GB300 NVL72, its current shipping platform.
For context, GB300 NVL72 was itself already 15 times more efficient than the generation before it, the Hopper architecture, on the same workloads.
| System | Throughput vs. previous gen | Token cost vs. GB300 NVL72 |
|---|---|---|
| Hopper (previous gen) | baseline | not measured |
| GB300 NVL72 | 15x better per megawatt | baseline |
| Vera Rubin NVL72 | 30x better per megawatt | 35x lower |
The tests used a benchmark called SemiAnalysis AgentX, which replays real-world AI coding sessions, complete with the tool calls and sub-tasks a genuine agent would trigger. NVIDIA notes that SemiAnalysis is still reviewing the results, and that the figures do not yet account for a key part of the chip stack, the Vera CPU, which handles tool-calling tasks.
What does this mean for ordinary people?
Most readers will not buy one of these systems. The impact arrives indirectly, through the products companies build on top.
When the cost of running an AI agent drops sharply, businesses can afford to let those agents work longer, on more customers, without the economics falling apart. That means the AI assistant inside a customer-service app, a health records system, or an accounting tool can handle harder problems rather than handing off to a human the moment a task gets complicated.
For the companies operating AI services, the electricity saving matters as much as the raw speed. Data centres face strict power limits, and 30 times more work per watt means fitting far more productive capacity into the same building.
Common questions
Are these numbers independently verified?
Not yet. NVIDIA measured these figures using the SemiAnalysis AgentX benchmark, and the results are currently pending review by SemiAnalysis. The company also notes the figures leave out Vera CPU performance, so final numbers could shift in either direction.
When can customers use Vera Rubin NVL72?
NVIDIA says Vera Rubin is in full production and scaling across its partner ecosystem now.



