Tech

Fish Audio secures $50 million seed to scale AI voice models for enterprise and creators

Led by Coreline Ventures and Capital Today, the funding round supports the development of advanced synthetic voice technology amidst a crowded market and ongoing debates over creator consent.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · original
Fish Audio raises $50M seed to build AI voice models for creators and enterprises
Palo Alto-based startup reports $21 million in annual recurring revenue as it pivots from open-source project to commercial platform

Fish Audio, a Palo Alto-based artificial intelligence voice startup, has raised $50 million in a seed funding round led by Coreline Ventures and Capital Today. The investment, which also includes participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0, is designated to accelerate the development of advanced AI voice models for both individual creators and enterprise clients.

Since launching its open-source and hosted platforms last year, the company has attracted more than 8 million users and generated $21 million in annual recurring revenue. The startup’s GitHub repository, Fish Speech, has accumulated over 31,000 stars, reflecting significant engagement from indie developers and video game designers. Founded by former NVIDIA researcher Shijia Liao, the technology originated as a small project to address the lack of expressive synthetic voices, initially trained on a single GPU.

The company has released five models in the past 12 months, comprising four speech generation models and one speech-to-text model. While three of its speech generation models remain open-source, its latest iteration, the S2.1 Pro model, is available exclusively through a paid API. Fish Audio offers tiered monthly plans for creators and teams, alongside an enterprise API version currently utilised by organisations such as HeyGen, Sanas, and Plaud.

Addressing past controversies regarding unconsented voice uploads, the startup has implemented an automated takedown process. CEO and co-founder Rissa Cao stated that creators can now submit voice samples or contracts to verify ownership, resulting in removal times of less than three minutes. However, industry partners note that the community-driven model requires robust mechanisms for verified voice ownership, clear licensing, and revenue-sharing to ensure long-term trust and durability.

Looking ahead, Fish Audio plans to release an audio understanding model this year and is developing a speech-to-speech model. The AI voice generation market remains highly competitive, with rivals including ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp. Investors believe that Fish Audio’s fine-grained controls and cost-efficient training capabilities will help it compete effectively against larger AI laboratories.

Continue reading

More from Tech

Read next: France Enacts Strict Ban on Unsolicited Telemarketing Calls
Read next: OpenAI expands Daybreak cybersecurity programme with new model tiers
Read next: AI models map 766 genes in schizophrenia genetic architecture