NVIDIA's Free, Open Voice AI Now Speaks 12 Languages and Responds in Under a Tenth of a Second

A new version of NVIDIA's Magpie text-to-speech model adds Arabic, Korean and Portuguese, and runs fast enough on your own hardware to keep voice assistants feeling natural.

AI2Day NewsdeskUpdated Editor: Lee Brown4 min read
A vast server room filled with rows of illuminated blue and white rack-mounted computers stretching into the distance, cool dramatic lighting reflecting off pol
Share

Key points

  • NVIDIA Magpie Multilingual TTS, a free open-weights text-to-speech model, now supports 12 languages after adding Arabic, Korean and Brazilian Portuguese in its latest release.
  • On NVIDIA's B200 chip, the model produces its first spoken audio in just 32 milliseconds, well inside the threshold where humans notice a delay.
  • Developers can download and run the model entirely on their own computers or servers, with no data sent to NVIDIA's cloud.
  • French character error rate, a measure of how accurately the model pronounces words, dropped from 2.70% to 1.54% compared with the previous version.
  • The model's architecture and benchmarks are documented in a paper accepted to ICASSP 2026, a leading audio research conference.

When you ask a voice assistant a question, a small clock starts ticking the moment you stop speaking. Your words get transcribed, an AI works out a reply, and then a text-to-speech engine turns that reply into sound. That last step is the one your ears judge most harshly. If the voice takes too long to start talking, the whole exchange feels broken.

NVIDIA wants to fix that with Magpie Multilingual TTS, a text-to-speech model released as open weights, meaning the full trained model files are free to download and run anywhere. The latest update, first detailed on Hugging Face, expands the model to 12 languages and improves naturalness across several existing ones. We first covered the original Magpie release on 10 August 2026, when on-device voice AI was also making moves at Apple.

What did NVIDIA actually change?

Three languages join the roster: Modern Standard Arabic, Korean and Brazilian Portuguese. That brings the total to 12, all handled by a single 364-million-parameter download rather than separate regional models.

Quality improved for existing languages too. French pronunciation tightened sharply; Spanish followed closely.

Language Error rate before Error rate now Voice match before Voice match now
French 2.70% 1.54% 0.703 0.747
Spanish 1.14% 0.60% 0.715 0.793
German 0.66% 0.80% 0.626 0.742

Source: Magpie TTS Multilingual model card. Lower error rate is better; higher voice match score is better.

How fast does it actually speak?

Fast enough that most people won't notice a pause. On an NVIDIA B200, the kind of chip found in high-end data centres, the model starts producing audio 32 milliseconds after receiving text. That's roughly a thirtieth of a second. On an A100, a widely deployed older chip, that delay is 79 milliseconds.

Humans begin to perceive a conversation as awkward when responses exceed roughly 200 milliseconds. Magpie's speech step leaves most of that budget free for transcription and inference.

Two engineering choices drive the speed. The model predicts two chunks of audio at once rather than one during each processing step, halving the iterations needed. A small additional network called a local transformer then cleans up any quality loss from doing so, keeping the voice sounding natural.

Why does running it yourself matter?

Most commercial text-to-speech services are cloud-based: your text travels to a remote server, gets converted to audio, and travels back. Every trip costs time and puts a third party in contact with your data.

Because Magpie's weights are open, a healthcare provider or financial institution with strict data-residency rules can run it on their own servers. Audio never leaves the building. Developers can also fine-tune the model on their own voice samples or specialist vocabulary, which a locked commercial API won't allow. That's the practical gap between this and something like Fish Audio's cloud service, where you hand your data to someone else's infrastructure.

The numbers that matter most here aren't the latency figures, impressive as 32 ms is. They're the quality scores: German's error rate actually crept up slightly in this release, from 0.66% to 0.80%, even as French and Spanish improved sharply. That's the one to watch as NVIDIA iterates toward a genuinely balanced multilingual model.

© 2026 AI2Day