NVIDIA's NVLink Fusion Lets Chipmakers Plug Their Own AI Chips Into NVIDIA's Data Centre Infrastructure

A new programme lets companies building their own specialised AI chips skip years of infrastructure work by connecting directly into NVIDIA's existing data centre plumbing.

AI2Day Newsdesk4 min read
Photoreal news-editorial style, 16:9 framing, edge-to-edge
Share

Key points

  • NVIDIA's NVLink Fusion programme lets companies plug their own custom AI chips, called XPUs, into NVIDIA's data centre networking and rack hardware.
  • Sixth-generation NVLink delivers end-to-end chip-to-chip transfer latency three times lower than standard Ethernet alternatives, with ten times the data packet rate.
  • Partners already confirmed include Intel, MediaTek, Amazon's Annapurna Labs, and manufacturing firm QCT.
  • NVLink Fusion supports domains of up to 72 chips today, with roadmap plans to scale to 1,152 accelerators in a single connected group.
  • The programme uses NVIDIA's MGX rack design, meaning custom-chip systems share the same physical shelving, cooling, and power setup as standard NVIDIA GPU servers.

Building a chip is only half the battle. The harder half is everything around it.

When a company designs a custom AI processor, it still needs high-speed connections between chips, cooling systems, power delivery, management software, and a supply chain for every bolt and bracket. That surrounding infrastructure can take years to build and test. NVIDIA's new NVLink Fusion programme, announced on the company's official blog, offers a shortcut: snap your chip into NVIDIA's already-proven setup and skip most of that work.

What exactly is NVLink Fusion?

NVLink Fusion is a programme that lets outside companies connect their custom processors, what NVIDIA calls XPUs (essentially any specialised AI chip that is not a standard NVIDIA GPU), into NVIDIA's NVLink network. NVLink is the high-speed cable technology NVIDIA uses internally to let its own chips talk to each other at very high speed.

Think of it like a universal power socket standard. Instead of every appliance manufacturer inventing its own wall socket, they all use the same format. Fusion gives custom chip designers a standard socket into one of the world's most widely deployed AI data centre platforms.

The performance numbers are striking. Chip-to-chip data transfers across NVLink run with latency, meaning delay, three times lower than connections built on standard off-the-shelf Ethernet cable. Packet rate, the volume of data chunks moving per second, is ten times higher. For AI workloads running huge models with trillions of parameters, that gap in speed directly affects cost: slower connections mean chips sit idle more often, and idle chips cost money without producing output.

Why does this matter for ordinary people?

Faster, cheaper AI infrastructure eventually means lower costs for every product or service running AI in the background, from customer service chatbots to medical image analysis tools. More competition in AI chip design also matters. Right now a small number of suppliers dominate the market for the hardware that runs AI. Fusion lets more companies build viable chips without building an entire data centre ecosystem from scratch, which could bring more suppliers into a market that badly needs them.

For the companies doing the building, the benefits are concrete. Intel's data centre chip team said the programme lets customers choose their preferred processor architecture without having to rebuild the infrastructure around it. Amazon's chip division, Annapurna Labs, noted it gains access to a proven rack design and a broader supplier base, both of which speed up delivery to customers.

What does the physical setup look like?

A Fusion-based server rack uses NVIDIA's MGX rack format. Compute trays, the slide-in modules holding the chips, use full liquid cooling with no fans, cables, or hoses inside the tray itself. Individual trays can be removed for repair while the rest of the rack keeps running. The networking switches connecting chips together are also liquid cooled and stay online during maintenance.

Software ties it together. NVIDIA supplies tools for distributing AI workloads across chips, managing the cluster, monitoring health, and debugging problems, all of which work whether the rack holds NVIDIA's own GPUs, custom XPUs, or a mixture of both.

QCT, a server manufacturing partner, said Vera Rubin NVL72 production already runs at close to 100 percent automated assembly, and that investment carries over directly to Fusion-based systems.

Common questions

Does this mean NVIDIA controls all AI chips now?

No. Fusion lets other companies use NVIDIA's networking and rack technology without giving up control of their own chip designs. The processor itself remains the partner's own product.

Will this change anything for regular users of AI tools?

Not immediately and not visibly. Over time, more chip competition and lower infrastructure costs tend to reduce the price of running AI services, which can filter through to cheaper or faster products.

© 2026 AI2Day