Suno Now Generates Spoken Voices, Not Just Songs

The AI music platform launches a public beta for speech synthesis, letting creators produce voiceovers and backing tracks in the same session. It's a real pivot for a company that, until two weeks ago, made only music.

AI2Day NewsdeskAI-assistedPublished Editor: Lee Brown3 min read
Illustration: a vintage studio microphone on a desk
Illustration made with AI. Not a photograph of the events described.
Share

Key points

Suno, the AI music platform best known for turning a text prompt into a full song, can now generate human-sounding spoken voices. The feature, called Speech, launched this week in public beta on its website and mobile apps, as first reported by The Verge AI.

What does this actually do?

Users type a script or describe the kind of voice they want, and Suno produces spoken audio. The key addition is that it works alongside the platform's music generation, so a creator can get a voiceover and a backing track in one session rather than stitching files together from different tools. Podcast intros and short-form video narration are the obvious uses. Whether the voice quality holds up against dedicated speech tools is what the beta period is meant to test.

Why does this matter for Suno's position?

This is a genuine change of direction. Until this week, Suno's entire product was music. The company has moved fast: Suno v6 became the first AI music model trained on songs licensed from record labels just two weeks ago, a significant legal shift after years of disputes with Sony, Universal and others.

Adding speech synthesis puts Suno in a new competitive bracket. It's no longer competing only with AI music tools; it now overlaps with voice platforms like ElevenLabs and the speech features built into general-purpose AI assistants.

The licensing question matters here too. Suno's music training raised enough copyright concerns that record labels sued before eventually licensing their catalogues for v6. Voice synthesis carries its own legal risks around consent and the use of real performers' voices as training data, and Suno hasn't yet said what data the speech model was trained on. We covered the broader pattern of legal pressure on Suno in our August report on watermarking and copyright-detection measures.

For ordinary users, the practical upside is simpler production. Generating a short explainer video or a demo reel no longer requires separate tools for voice and music. That convenience is real, and it's probably what most people will notice first.

Common questions

Is the speech feature free to use?

Suno hasn't announced separate pricing for Speech. As a public beta, it appears within the existing platform tiers, but the company may change access terms before a full launch.

Can the tool clone a specific person's voice?

Suno's announcement describes generating voices from text descriptions or scripts, not from uploading a sample of an existing voice. Voice cloning doesn't appear to be part of this feature.

Who is affected first?

Anyone with a Suno account on web or mobile can access the beta now. Professional creators producing short-form video or podcast content stand to benefit most from the combined voice-and-music workflow.

© 2026 AI2Day