Fish Audio provides text-to-speech, voice cloning, speech-to-text, voice agents, and other audio tools. Its TTS controls include emotion and expression, real-time generation, multilingual support, and adjustments for speed, volume, and model parameters. Fish Audio states that TTS supports eight languages with native accents and that its platform hosts more than 2,000,000 voices. Developers can use the API for speech generation, voice cloning, and transcription through REST, WebSocket streaming, and official Python and TypeScript SDKs. Listed platforms include web, Linux, macOS, Windows, API, and self-hosted deployments. The free tier includes 8,000 monthly credits, up to seven minutes of generation, and personal, non-commercial use. Commercial use is listed as available, and paid plans start at $15/mo. Enterprise deployments include VPC, on-premises, air-gapped, and sovereign cloud options. Fish Audio says default data stays in the United States and self-hosted deployments run inside the customer's infrastructure. Unused monthly minutes do not roll over.
Who it is for
Fish Audio may suit creators producing voiceovers, audiobook narration, or character voices, as well as developers building conversational chatbots. Its API, SDKs, and enterprise deployment options also serve technical teams.
What is good
- TTS includes emotion, expression, and speed controls.
- Supports eight languages with native accents.
- API supports generation, cloning, and transcription.
- Commercial use is listed as available.
What to know first
- Free tier is for personal, non-commercial use.
- Unused monthly minutes do not roll over.
- SOC 2 Type II audit is underway.
Freedom251 review
Fish Audio: the full review
Fish Audio combines audio generation tools with voice cloning, transcription, and developer access. Note the free tier’s usage and commercial-use limits, and that monthly minutes expire at the billing-cycle boundary.
Fish Audio is an audio-generation service for creators and developers who need speech synthesis alongside cloning, transcription, or voice-agent tools. It is most compelling when adjustable delivery and API access matter together, but the free tier is strictly non-commercial and the larger minute allowances require annual billing.
Overview
Fish Audio brings text-to-speech, voice cloning, speech-to-text, and voice agents into one product. Its use cases range from video voiceovers, audiobook narration, and character voices to conversational chatbots. That breadth suits people moving between content production and audio software development; it is less compelling if you need only a narrow, single-purpose tool.
The TTS system offers emotion and expression controls, real-time generation, and adjustments for speed, volume, and model parameters. It automatically supports eight languages with native accents. A library of more than 2,000,000 voices and output in WAV, PCM, MP3, or Opus offer substantial choice, though a large catalog is most useful when its voices fit the production at hand.
Key features
- Generation and cloning: Voice cloning sits alongside TTS controls for delivery, giving creators more than basic read-aloud output. The Free Tier's 500-character generation ceiling makes it a poor fit for long passages; paid plans allow longer individual generations.
- Developer access: The API covers speech generation, cloning, and transcription through REST and WebSocket streaming, with official Python and TypeScript SDKs. That combination is relevant for both request-based integrations and streaming applications.
- Team and enterprise use: Integrations or ecosystem connections include Vapi, Twilio, Retell, and workflow automation tools. Enterprise deployments can run in VPC, on-premises, air-gapped, or sovereign cloud environments; self-hosted deployments run inside the customer's infrastructure.
- Data and support: Default data stays in the United States. Fish Audio says its SOC 2 Type II audit is underway; enterprise contracts can enable Zero Data Retention, and it describes HIPAA-aligned configurations and BAAs for qualifying healthcare workloads. Enterprise support includes 24/7 production support, a technical account manager, and a stated 99% uptime SLA.
Pricing
Fish Audio is freemium. Its Free Tier costs 0.00 USD per free and provides 8,000 monthly credits, up to 7 minutes of generation, up to 500 characters per generation, three public voice slots, and standard generation speed. It is for personal, non-commercial use only, so anyone publishing or monetizing the output should look to a paid plan.
Plus costs 11.00 USD per month, billed $132 billed annually. It includes 250,000 monthly credits, up to 200 minutes, 15,000 characters per generation, unlimited public and 10 private voice slots, and one professional voice slot. It is a practical step up for individual creators who need longer generations and private voices, but it does not include team seats in the stated plan terms.
Pro costs 75.00 USD per month, billed $900 billed annually. Its 2,000,000 monthly credits support up to 1,620 minutes and 30,000 characters per generation; the plan also includes three team seats, unlimited voice slots, five professional voice slots, and a seven-day money-back guarantee. It fits collaborative production better than Plus, though its annual bill is a substantial commitment.
Max costs 749.00 USD per month, billed $8988 billed annually. It raises the allowance to 25,000,000 monthly credits and up to 6,250 minutes, with 10 team seats and 15 professional voice slots. That scale is aimed at high-volume teams, not occasional generation.
Enterprise has custom pricing and offers volume pricing, pay-as-you-go organization controls, Zero Data Retention, on-premises deployment, and SOC2 compliance. Across the monthly-minute plans, unused minutes do not roll into the next billing cycle, so demand that varies month to month can leave paid capacity unused.
Platforms
Fish Audio supports API, Linux, macOS, self-hosted, web, and Windows. This range gives developers and organizations deployment options beyond a browser-based workflow.
Who it's for
Choose Fish Audio if you create voiceovers, narration, or character audio and want control over delivery, or if you are building an application that needs generation, cloning, or transcription through an API. Teams with deployment or data-handling requirements may find the enterprise options relevant. It is a weaker fit for commercial work on a no-cost plan or for buyers who need monthly minutes to carry forward.
Pros and cons
- Pros: TTS controls, real-time generation, and eight supported languages provide meaningful flexibility in voice output.
- Pros: REST, WebSocket, and official Python and TypeScript SDKs cover several common development approaches.
- Pros: Enterprise deployment choices and optional Zero Data Retention address needs that a consumer-only service may not serve.
- Cons: The free plan prohibits commercial use and caps generation at seven minutes monthly and 500 characters per generation.
- Cons: Monthly minutes expire at the billing-cycle boundary, which disadvantages users with uneven workloads.
- Cons: The paid plans are quoted with annual billing totals, making the commitment clearer than a flexible month-to-month option.
Alternatives
AI Voice Cloning Software, Voice Cloning Software, AI Voice Generators, and Text-to-Speech Software are category lists for comparing options by workflow.
- GPT-SoVITS is a free option with API, desktop, web, and self-hosted platforms if those deployment choices matter more than Fish Audio's published paid tiers.
- Voice.ai has a free plan with 5k monthly credits, an online voice changer, text to speech, and audio tools with a five-minute conversion limit; consider it if that free tool mix is a closer fit.
- Voicebox is another free-plan option, with Android, iOS, Linux, macOS, Windows, API, and self-hosted platforms.
- VoiceStudio is free and supports API, desktop, web, and self-hosted platforms, making it an alternative for readers focused on those access options.
- Altered Studio has a free trial and a free plan with three minutes of monthly voice morphing, local voice cloning, and noncommercial use; consider it for that limited free workflow.
- Speechify offers a free plan with 10 voices and a 1.5x maximum speed, and may suit readers prioritizing those terms.
- Applio is free forever, open source under MIT, and requires no accounts or paywalls; choose it if those terms and its Linux, macOS, self-hosted, web, or Windows platforms are the priority.
- iMyFone VoxBox is another freemium option, with web, desktop, and mobile platforms.
Verdict
Fish Audio is a strong choice for creators and developers who want controllable speech generation, cloning, and transcription under one roof, particularly when API or enterprise deployment options matter. Look elsewhere if commercial use must be free, if you need monthly minutes to roll over, or if the annual commitments attached to its paid tiers do not match your workload.
Fish Audio plans and pricing
All plansCompared on text-to-speech software
- Free plan
- Yes
- Paid from
- $15/mo
- API access
- Yes
- Commercial use
- Yes





