Voicebox is an open-source local voice studio for voice cloning, speech generation, dictation, and agent voice output. Its desktop app can clone a voice from a brief audio sample provided by upload, microphone, or system capture. The maker lists seven text-to-speech engines, plus a timeline editor for multi-voice stories, track arrangement, trimming, and conversation mixing. Effects include pitch shift, reverb, delay, and compression. A keyboard shortcut can send dictation to the focused text field or clipboard; local Whisper models support 99 languages. The app works offline without an account, and its privacy policy says it does not transmit recordings, profiles, transcripts, generations, or settings to the maker. A local REST API supports speech generation and related functions, and compatible agents can speak through Voicebox. The Local plan is free forever, with unlimited local generations and captures. Cloud backup and mobile sync are marked as coming soon, and the listed cloud pricing and limits are not final.
Who it is for
Voicebox suits people who want local voice generation, cloning, dictation, or agent voice output. Its listed uses include game dialogue, accessibility readouts, audiobooks, podcast intros, apps, and scripts or tools.
What is good
- Works offline without an account.
- Local plan includes unlimited generations and captures.
- Local Whisper models support 99 languages.
- Local REST API supports speech generation.
What to know first
- Cloud backup and mobile sync are coming soon.
- Cloud pricing and limits are not final.
- Desktop downloads list specific macOS and Windows builds; Linux is built from source.
Freedom251 review
Voicebox: the full review
Voicebox centers voice work on a local desktop app, with generation, dictation, editing, and API features. The Local plan is free forever; cloud features are still marked as coming soon.
Voicebox is a local-first voice studio for people who want to create, edit, or route speech from a desktop. It is strongest for creators and developers who value offline processing and an open-source app; readers who need mobile or cloud workflows should wait or choose another tool.
Overview
Voicebox brings voice cloning, speech generation, dictation, audio editing, and agent voice output into one desktop app. A clone can begin with a three-second audio sample, supplied by upload, microphone recording, or system audio capture. Support for dubbing and 23 languages broadens its reach beyond single-language narration.
The local app works offline without an account. Its privacy policy says recordings, voice profiles, transcripts, generations, and settings are not sent to the maker. That makes Voicebox a considered option for work users prefer to keep on their own machine, though cloud sync and mobile use are not ready yet.
Voicebox fits Voice Cloning Software, AI Voice Cloning Software, and Text-to-Speech Software workflows. Its pitch and other sound effects also place it among Voice Changer Software.
Key features
Generation, cloning, and stories
Seven text-to-speech engines cover cloning and preset voices, with delivery controls. For creators, that range and the short source-audio requirement make it practical to move between original voices and generated speech. The timeline-based Stories editor supports multiple voices, track arrangement, clip trimming, and conversation mixing, giving longer narrative projects more structure than a basic text-to-speech interface.
Effects and dictation
Pitch shift, reverb, delay, compression, and other effects can be previewed live and saved as presets. That is useful when shaping a voice inside a project, though the app’s broader editing focus may be unnecessary for someone who only needs straightforward read-aloud output. Dictation works through a global keyboard shortcut, sending a transcript to the focused text field or clipboard. Local Whisper models support 99 languages, a separate and broader language count than the 23 languages listed for Voicebox overall.
Agent and API output
MCP-aware agents such as Claude Code, Cursor, and Cline can speak through Voicebox using a cloned voice. A local REST API covers speech generation, voice profiles, model status, generation history, and health checks; the maker also describes a POST /speak endpoint for other clients. These integrations make it a stronger fit for developers building voice into tools than for users seeking only a standalone narrator.
Pricing
Local: 0.00 USD per free, billed Free forever. It includes the full app, unlimited local generations and captures, open source, and no account requirement. This is the clearest choice for individuals who can work on one desktop and do not need cloud backup.
Cloud: 12.00 USD per year, billed $12/year at a launch price. The planned tier includes 25 GB of encrypted storage, up to five devices, 30-day version history, and encrypted backup and sync. Cloud is coming soon, and its pricing and limits are not final, so this is not yet a dependable option for buyers who need sync today.
Studio: 48.00 USD per year, billed $48/year as a placeholder with pricing TBD. It is planned to include 250 GB storage, unlimited devices, one-year version history, and priority support. This tier is aimed at heavier multi-device use, but its provisional price and future status make it difficult to compare as a current purchase.
There is no free trial. The free Local plan has no stated generation or capture cap, while the paid tiers add storage, device allowances, history, and backup rather than more local generation capacity.
Platforms
Downloads are listed for macOS on Apple Silicon and Intel, Windows 64-bit, and Linux builds from source. The broader platform list also names Android, iOS, API, and self-hosting, but mobile and cloud sync are described as coming soon. Docker deployment is documented. Commercial use and API access are supported, making the local desktop useful for commercial workflows as well as personal projects.
Who it's for
Voicebox is best for desktop creators producing game dialogue, audiobooks, podcast intros, or scripts; developers adding voice output to apps and agents; and accessibility users who want dictation or readouts without an online account. It is less suitable for teams that need a ready-made mobile workflow, finished cloud synchronization, or confirmed paid-plan terms.
Pros and cons
- Pros: The free plan includes the full app and unlimited local generations and captures, with no account required.
- Pros: Cloning, multi-voice editing, effects, dictation, and agent/API output cover several voice workflows in one local tool.
- Pros: The privacy policy says user audio and related work stay off the maker’s servers.
- Cons: Mobile and cloud features remain future offerings, limiting use across devices today.
- Cons: Linux users must build from source, unlike the listed direct macOS and Windows downloads.
- Cons: Cloud and Studio terms are provisional, so their listed storage and device benefits are not a settled basis for purchase.
Alternatives
Voice.ai is worth considering for users who want a freemium voice changer and text-to-speech service across mobile, web, desktop, and API platforms; its Free plan has 5k credits/month.
Fish Audio is a better fit for users who prefer a web-accessible option alongside desktop, API, or self-hosted platforms; its Free Tier provides 8,000 monthly credits, up to seven minutes of generation, and up to 500 characters per generation.
Speechify suits readers prioritizing broad language support and a large voice selection: Premium is 29.00 USD per month and includes 60+ languages and 1,000+ voices, while its free plan offers 10 voices and a 1.5x maximum speed.
TTSReader is an option for straightforward text-to-speech across web and mobile, but its free plan limits premium-voice use to 5,000 characters and excludes commercial use, API, and automation.
Typecast is an alternative for users who want video exports alongside voice work; its free plan includes three projects, 3,000 lifetime credits, 1 GB media storage, and 720p exports.
VoiceStudio is another free option.
Altered Studio may suit users seeking voice morphing with a limited free monthly allowance and a noncommercial attribution license.
Audexum Text to Speech is another freemium alternative.
Verdict
Choose Voicebox if you want a free, open-source desktop studio that combines local voice creation with editing, dictation, and developer integrations. Its strongest reason to buy in is the breadth of offline work without an account; its clearest reason to look elsewhere is that cloud and mobile workflows are still forthcoming.
Voicebox plans and pricing
All plansCompared on text-to-speech software
- Free plan
- Yes
- Cloning method
- instant
- Supported languages
- 23 languages
- Dubbing workflow
- Yes
- API access
- Yes
- Commercial use
- Yes





