Nvidia's $20 Billion Groq Bet Goes Live: What It Means for AI Speed
Nvidia says its Groq 3 LPX racks are now in full production and will be online later this year, promising dramatically faster AI responses for everyday users.

Key points
- Nvidia confirmed its Groq 3 LPX rack is in full production as of March 2026, following its $20 billion acquisition of Groq's assets in December 2024.
- Each LPX rack packs 256 individual Groq 3 chips and can deliver 3,400 tokens per second, a measure of how fast an AI model spits out text.
- The racks will first go live at Nebius, a cloud computing provider, alongside Nvidia's new Vera central processors and Rubin graphics processors.
- Nvidia CEO Jensen Huang has projected $1 trillion in combined sales for Blackwell and Vera Rubin chips through 2027.
- Rival chip efforts from Advanced Micro Devices and Cerebras are chasing the same goal: faster, cheaper AI responses at scale.
You know that spinning circle you sometimes see when asking an AI chatbot a question? That delay has a name: latency. And right now, cutting it is the hottest race in AI hardware.
Nvidia announced Monday that its Groq 3 LPX rack, a large server unit containing 256 specialised chips, is now in full production and will go online later this year. CNBC Tech first reported the details from an Nvidia briefing with reporters.
What is this chip actually for?
Think of it as a speed specialist. Standard GPUs, the powerful chips that do most of AI's heavy lifting, are flexible workhorses: they can train AI models and run them. Groq chips do one thing extremely well: they make already-trained AI models respond faster.
Specifically, they speed up what engineers call the "decode" phase, the part where an AI model generates each word of its reply. The Groq 3 chip keeps 500 megabytes of very fast memory, called SRAM, directly on the chip itself. That cuts out the usual back-and-forth to slower external memory, which is where a lot of delay creeps in.
The result: Nvidia says its Groq 3 LPX rack hits 3,400 tokens per second (tokens are roughly chunks of words, so this translates to extremely fast text generation), based on a benchmark from research firm Artificial Analysis.
Why does this matter to ordinary users?
Faster responses make AI agents, software that can carry out multi-step tasks on your behalf like booking travel or writing code, feel snappy rather than sluggish. Coding assistants in particular demand this kind of speed.
For cloud companies, it also opens a business opportunity: charging a premium for guaranteed fast responses. Nvidia senior director Dion Harris put it plainly: faster chips "unlock the ability to offer premium tiers of service" for customers who cannot afford to wait.
Nvidia's Groq chips are made by Samsung, while its standard GPUs come from TSMC (Taiwan Semiconductor Manufacturing Company, the world's largest contract chipmaker). The racks will be deployed at Nebius, a neocloud provider, sitting alongside Nvidia's Vera and Rubin systems.
How does it stack up against the competition?
| Product | Tokens per second | Status |
|---|---|---|
| Nvidia Groq 3 LPX rack | 3,400 | In production, online later 2026 |
| OpenAI Ultrafast mode | 750 | Available now (powered by Cerebras) |
| AMD + Cerebras rack integration | Not disclosed | Announced 2026 |
OpenAI's Ultrafast mode, announced recently and powered by chips from Cerebras (a rival AI chip company that recently went public on the stock market), currently promises 750 tokens per second. Nvidia's claimed 3,400 is a significant step beyond that, though real-world results often differ from benchmarks.
Nvidia CEO Jensen Huang said at the Groq 3 launch in March that he would dedicate roughly a quarter of data centre space aimed at coding tasks to Groq chips. The rest, he said, stays with Vera Rubin.
Nvidia reports its quarterly earnings on Wednesday, where investors will be watching closely for any update on how quickly these new systems are selling.



