The demand for AI-generated voice technology is growing rapidly across two distinct fronts. On one side, creative professionals need voice models capable of nuanced, expressive delivery. On the other, businesses looking to automate customer-facing operations such as support lines and sales processes require voices that are precise, controllable, and consistent. One startup is positioning itself to serve both camps simultaneously.
Fish Audio, headquartered in Palo Alto, has built a library of more than 15,000 natural language controls designed to cover a wide spectrum of voice generation use cases. Since its launch last year, the platform has attracted over 8 million users across its open source and hosted offerings, and the company now reports an annual recurring revenue figure of $21 million.
On Tuesday, Fish Audio announced it has secured $52 million in a seed funding round. The round was led by Coreline Ventures and Capital Today, with additional participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The company intends to use the capital to develop more sophisticated models and expand its enterprise-grade offerings.
From a Single GPU to 31,000 GitHub Stars
Fish Audio originated as a personal side project by Shijia Liao, a former researcher at Nvidia. Dissatisfied with the flat, robotic quality of synthetic voices available at the time, Liao trained a voice-generation model on a single GPU and subsequently released it as an open source project. That repository, known as Fish Speech on GitHub, has since accumulated more than 31,000 stars and found a user base that includes independent software developers, video game designers, and content creators.
Over the past year, the company has released five models in total - four focused on speech generation and one dedicated to speech-to-text conversion. Three of its speech-generation models remain open source, while its most recent release, the S2.1 Pro model, is available exclusively through a paid API. This tiered approach allows the company to maintain its open source community while building a commercial revenue stream around its most advanced capabilities.
A Platform Built for Creators and Enterprises Alike
Fish Audio offers monthly subscription plans tailored to individual creators and teams. These plans provide access to a set allocation of generation minutes along with voice-cloning features. For larger organizations, the company offers enterprise-level API access and a dedicated platform. Notable clients already using the platform include HeyGen, a company that powers AI-driven video avatars, and Sanas, which focuses on real-time accent and speech modulation technology.
Rissa Cao, CEO and co-founder of Fish Audio, described how the requirements differ significantly depending on the type of customer the platform serves.
"Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls."
Voice Consent and Takedown Concerns
Part of how Fish Audio has built its voice library is through a user contribution model, in which people submit their own voice recordings for use in model training and receive compensation if their voices are subsequently used. However, this approach drew criticism earlier this year when a number of creators alleged that their voices had been uploaded to the platform without their knowledge or approval.
While Fish Audio had a DMCA takedown mechanism in place, complaints emerged that the process was slow and cumbersome. Cao said the company has since automated the procedure. Creators can now submit either a short voice sample or a formal contract as proof of ownership, and the company says the voice in question will be removed from the platform within three minutes of a verified claim.
That said, the underlying issue remains unresolved in a broader sense. There is currently nothing stopping a third party from uploading someone else's voice without their awareness. Until the affected creator discovers the unauthorized upload and files for removal, their voice remains available on the platform.
Osuke Honda, a partner at Coreline Ventures, acknowledged that the community-based model carries inherent trust obligations that must be treated as core product features rather than compliance checkboxes.
"A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially."
Why the Startup Decided to Raise Outside Capital
Cao explained that for much of its early life, Fish Audio operated efficiently without external investment. As long as the company was focused purely on its open source project and creator-facing subscription plans, it did not feel the need for venture backing. However, as the company's ambitions grew to include more advanced model development and enterprise-scale infrastructure - and as investor interest intensified - the decision was made to pursue outside funding.
Looking ahead, Fish Audio has outlined plans to release an audio understanding model before the end of the year. The company is also working on a speech-to-speech model, which would allow real-time voice transformation and interaction rather than simply generating speech from text inputs.
A Crowded Market With Clear Competition
Fish Audio is entering a well-established competitive landscape. The speech synthesis and voice generation market already includes prominent players such as ElevenLabs, WellSaid Labs, Cartesia, Speechify, Async (formerly known as Podcastle), and Krisp, all of which are competing for budget allocations from both individual creators and enterprise clients.
Rico Mallozzi, a partner at 359 Capital, argued that Fish Audio's ability to offer developers fine-grained controls combined with cost-efficient model training gives it a meaningful edge relative to larger and better-funded AI laboratories.
"I think what they've been able to build - state-of-the-art models - with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices."
With $52 million now secured, Fish Audio is positioned to accelerate its model development roadmap, deepen its enterprise relationships, and address the trust and consent challenges that have surfaced as its platform has scaled. How it navigates those issues - particularly around voice ownership and creator rights - may prove just as consequential as the technical progress it makes in the months ahead.



