Fish Audio raises $50 million to let anyone clone a voice, but consent questions linger
The Palo Alto startup has 8 million users and $21 million in annual revenue. Its new funding will push into enterprise contracts and more advanced models. A recent controversy over unauthorised voice uploads is not fully resolved.

Key points
- Fish Audio closed a $50 million seed round on Tuesday, led by Coreline Ventures and Capital Today.
- The startup reported $21 million in annual recurring revenue and more than 8 million users as of this week.
- Fish Audio's voice library was built partly from user-submitted recordings, some of which were uploaded without the original creators' consent.
- The company automated its voice take-down process, promising removal in under three minutes after a verified claim.
- Fish Audio plans to release an audio-understanding model and a speech-to-speech model before the end of 2025.
A Palo Alto startup that lets developers, game designers and businesses generate realistic synthetic voices has just secured a large injection of early-stage funding, confirming strong investor appetite for AI voice technology even as the sector wrestles with thorny questions about whose voice belongs to whom.
Fish Audio announced Tuesday that it raised $50 million in a seed round, the type of early funding that typically comes before a company has proven itself at scale. Coreline Ventures and Capital Today led the round, with 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0 also participating.
What does Fish Audio actually do?
The company builds text-to-speech models, software that converts written words into spoken audio, with a focus on expressiveness and fine-grained control. Its library offers more than 15,000 natural-language controls, meaning a developer can write an instruction like "sound nervous but reassuring" rather than adjusting technical sliders.
Fish Audio launched five models in the past year: four speech-generation models and one speech-to-text model, which does the reverse and converts audio into written text. Three of the speech models are open-source, meaning anyone can download and use the code freely. Its latest model, S2.1 Pro, is available only through a paid API, a programming interface that lets other apps plug into Fish Audio's technology.
The company traces its roots to a side project by former Nvidia researcher Shijia Liao. Frustrated by how robotic synthetic voices sounded, he trained a model on a single GPU, the specialised chip that does the heavy number-crunching AI requires, and posted the code publicly. That repository has since collected more than 31,000 stars on GitHub, a rough measure of developer interest.
Today, paying customers include HeyGen, which powers AI avatars, and enterprise voice-agent platforms. Fish Audio generates $21 million in annual recurring revenue, first reported by TechCrunch.
Should creators be worried about their voices being used without permission?
Possibly, yes. Earlier this year, some artists alleged their voices appeared on Fish Audio's platform without their consent. Part of Fish Audio's growth strategy involves asking users to submit voices for model training, compensating contributors when their recordings are used. The system relies on trust, and that trust cracked.
The company has since automated its DMCA take-down process. DMCA, the Digital Millennium Copyright Act, is a US law that gives creators a formal route to demand removal of their copyrighted material from online platforms. CEO Rissa Cao says a verified complaint now triggers removal in less than three minutes.
The gap in the process remains real, though. Nothing stops someone from uploading another person's voice without their knowledge. The voice stays live until the affected creator discovers it and files a complaint.
Oskue Honda, a partner at lead investor Coreline Ventures, acknowledged the problem directly: "Consent, transparency, and attribution must be built into the product rather than treated as afterthoughts."
What happens next?
Fish Audio will use the funding to develop more advanced models and court larger enterprise clients. Two new products are on the roadmap for 2025: an audio-understanding model, which would interpret the meaning and content of audio clips, and a speech-to-speech model that converts one spoken voice into another in real time.
The market is crowded. ElevenLabs, WellSaid, Cartesia and Speechify all compete for the same developers and business budgets. Fish Audio's argument is that it can match state-of-the-art quality at lower cost, partly because its lean founding team has kept training expenses down.
For ordinary users, the practical takeaway is straightforward: search your name or voice on any AI audio platform you have not explicitly signed up for. If you find something, most platforms now have a take-down route. Fish Audio's is now faster than most.
Common questions
Can a company legally use your voice to train an AI without asking you?
In most jurisdictions, no. Copyright law and, in some US states, right-of-publicity statutes give individuals control over commercial use of their voice. The legal picture is still evolving, but uploading someone else's voice without consent already carries real risk for the uploader.
What is a seed round, and why does $50 million sound large for one?
A seed round is the first formal outside investment a startup takes, usually used to build product and hire staff. Fifty million dollars is unusually large at this stage, reflecting how much competition there is to back AI infrastructure companies before they reach later, more expensive funding rounds.



