Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Hume launched Octave on February 26, 2025, as a text-to-speech model designed to use context and delivery instructions—not just pronunciation rules—to generate expressive speech. It can create voices from natural-language descriptions, clone a voice from a short recording, and adjust pacing, emphasis, pitch, tone, and emotional character.

The original launch is now only part of the story. Hume’s newer Octave 2 is documented as a live preview with broader language support, lower stated model latency, voice conversion, phoneme editing, and timestamp support. Because Octave 2 remains a preview, its capabilities, pricing, and availability may change.

What is Hume Octave?

Octave is Hume’s expressive text-to-speech system. Hume expands the name as “Omni-capable Text and Voice Engine” and describes it as a speech-language model: a system intended to model both language and speech rather than simply converting written characters into phonemes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In practical terms, Octave uses semantic context to influence how a line is spoken. The same sentence may be delivered differently depending on whether its context suggests excitement, disappointment, sarcasm, calm reassurance, or urgency. Hume’s documentation says the system can adapt pronunciation, pitch, tempo, and emphasis according to an utterance’s intended meaning.

#1 Best Overall
Plaud Note Pro AI Voice Recorder Transcribe & Summarize for Meetings Calls
  • ENHANCED CONTEXT WITH MULTIMODAL INPUT: Capture audio, type notes, add images, and press to highlight key moments for richer context. During recording, instantly mark key moments with a single button press. Simultaneously enrich your audio by snapping photos of important documents or typing in ideas
  • CHAT WITH YOUR RECORDINGS USING "ASK Plaud": Unlock deeper insights with this interactive AI. Ask questions, extract key points, draft emails, and get next-step suggestions—all grounded in your original audio for reliable, ready-to-use answers
  • INTELLIGENT RECORDING WITH AI DIRECTIONAL AUDIO: Enjoy seamless, intelligent recording with Plaud Note Pro. Its AI automatically switches between call and meeting modes while recording, while directional audio and real-time spatial awareness minimize noise to capture voices with crystal clarity
  • Everything Included: Includes Plaud Note Pro, magnetic case, magnetic ring, charging cable, and a free Starter Plan with 300 transcription minutes per month. Upgrade anytime in the Plaud app to Pro Plan (1,200 min/mo) or Unlimited Plan(Up to 24 hours of transcription per user per day)
  • PREMIUM ULTRA-SLIM DESIGN WITH INSTANTVIEW DISPLAY: Meticulously designed, the AI Note Taker is just 0.12 inches thin and 1.06 oz —about the size of a credit card. Its sleek aluminum body with a textured wave finish features a vivid AMOLED display, letting you check battery and recording status at a glance, while it seamlessly works with Apple Find My to ensure you never misplace it

That does not mean Octave experiences emotions or understands them like a person. “Understands” is best read operationally: the model uses meaning and context in the input to shape generated prosody and vocal delivery.

What launched in February 2025?

Hume’s February 26, 2025 announcement presented Octave as a model that could:

  • Generate a voice from a natural-language description.
  • Clone a voice from a short recording.
  • Perform character-style or “acting” deliveries.
  • Follow instructions that alter emotion, pacing, emphasis, and speaking style.
  • Use context-aware prosody in narration and interactive experiences.
  • Support different voices or personalities within an application.

Hume had introduced Octave in an earlier December 2024 announcement, but the February post framed the product as available through its platform and API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read Hume’s February 2025 launch announcement.

Voice design through ordinary language

Instead of choosing only from fixed acoustic settings, developers can describe the voice they want. Hume’s current voice documentation says its Voice Library contains more than 100 Hume-created voices, while custom voices can be designed through prompts. These voices can be used with Hume’s TTS and EVI products.

Useful descriptions can specify perceived age, accent, tone, personality, energy, emotional character, and speaking style. For example:

  • “A patient, empathetic counselor with a warm, measured delivery.”
  • “A rapid-fire Brooklyn cab driver with a nasal, high-energy voice.”
  • “A dramatic medieval knight speaking with restrained authority.”

These descriptions are guidance, not guaranteed controls over every acoustic property. Results can vary with wording, model version, language, and script complexity. A voice prompt that works well for a short line may not produce the same level of consistency across a long audiobook or game script.

See Hume’s voice design and cloning documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Plaud Note Pro AI Voice Recorder Transcribe & Summarize for Meetings Calls
  • AI-POWERED TRANSCRIPTION & SUMMARIES: Plaud Note Pro is your professional voice transcriber, delivering high-accuracy transcription in 112 languages with auto speaker labels. Powered by top AI models and thousands of templates, Note Pro instantly creates structured summaries, mind maps, To-Do lists, and proposals tailored to your role and industry
  • ENHANCED CONTEXT WITH MULTIMODAL INPUT: Capture audio, type notes, add images, and press to highlight key moments for richer context. During recording, instantly mark key moments with a single button press. Simultaneously enrich your audio by snapping photos of important documents or typing in ideas
  • CHAT WITH YOUR RECORDINGS USING "ASK Plaud": Unlock deeper insights with this interactive AI. Ask questions, extract key points, draft emails, and get next-step suggestions—all grounded in your original audio for reliable, ready-to-use answers
  • INTELLIGENT RECORDING WITH AI DIRECTIONAL AUDIO: Enjoy seamless, intelligent recording with Plaud Note Pro. Its AI automatically switches between call and meeting modes while recording, while directional audio and real-time spatial awareness minimize noise to capture voices with crystal clarity
  • Everything Included: Includes Plaud Note Pro, magnetic case, magnetic ring, charging cable, and a free Starter Plan with 300 transcription minutes per month. Upgrade anytime in the Plaud app to Pro Plan (1,200 min/mo) or Unlimited Plan(Up to 24 hours of transcription per user per day)

How Octave controls emotion and delivery

Octave has two important control layers.

1. Context in the spoken text

The words themselves provide clues about intent. “I can’t believe you actually came” could sound joyful, angry, disbelieving, or affectionate depending on the surrounding context.

2. Explicit performance instructions

Hume’s TTS API supports an utterance containing spoken text, an optional description, an optional voice, speed, and trailing silence. The spoken words belong in the text field; performance guidance belongs in description where supported.

Specific instructions are generally more useful than a single abstract label. A practical direction might describe intensity, pacing, pauses, audience, and how the delivery should change during the sentence:

“Deliver this with surprised delight, pause briefly after ‘believe,’ then soften at the end.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Teams should generate several versions rather than assuming one instruction will produce a perfectly repeatable performance. Expressiveness is not the same as deterministic control over every pause, syllable, or emotional transition.

Hume’s launch comparison

Hume reported a blind comparison involving 180 human raters and 120 diverse prompts against ElevenLabs Voice Design. In Hume’s results, Octave was preferred for:

Category Hume-reported preference
Audio quality 71.6%
Naturalness 51.7%
Matching the requested voice description 57.7%

These are Hume’s own study results, not an independent industry benchmark. The comparison was against a specific ElevenLabs feature, not every ElevenLabs model or product. The percentages are preference results rather than universal objective scores, so production teams should evaluate their own scripts, languages, voices, and delivery requirements.

Rank #3
Pocket AI Voice Recorder, Auto Transcription, AI Note Taker, Space Grey
  • YOUR AI PERSONAL ASSISTANT FOR EVERYDAY PRODUCTIVITY: More than a voice recorder, Pocket works as your AI personal assistant to capture, transcribe, and summarize meetings, calls, and ideas instantly. Core features are included out of the box, with optional advanced tools available for power users.
  • ONE-TAP RECORDING FOR REAL-LIFE MOMENTS: Capture meetings, phone calls, and in-person conversations instantly with a simple tap, no typing, no interruptions, just effortless note-taking anywhere you go.
  • SMART AI INSIGHTS & ORGANIZATION: Pocket automatically turns recordings into clear summaries, key action items and structured conversation maps so you can quickly review what matters without digging through audio.
  • TURN CONVERSATIONS INTO ACTION WITH “ASK POCKET”: Don’t just record, understand. Instantly ask questions across your meetings, extract key insights and generate next steps in seconds. All grounded in your recordings, so answers stay accurate and reliable.
  • MAGSAFE COMPATIBLE FOR SEAMLESS USE: Easily attach Pocket to your iPhone or other MagSafe compatible devices for convenient, hands-free recording on the go. Perfect for capturing meetings, calls, and ideas without needing to hold your device.

Review Hume’s reported comparison.

Octave 1 versus Octave 2 preview

Hume announced Octave 2 on October 1, 2025. Current documentation labels it a preview and distinguishes it from the original Octave model.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Capability Octave 1 Octave 2 preview
Languages English and Spanish Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian, and Spanish
Stated model latency Approximately 200 ms in the current comparison Approximately 100 ms, excluding network transit
Voice cloning Supported Supported; Hume advertises cloning from as little as 15 seconds of audio
Voice design Supported Current feature table lists voice design as English-only
Voice conversion Not established in the original launch material Documented for Octave 2
Word and phoneme timestamps Availability varies Supported when requested through the appropriate Octave 2 version
Status Original model Preview

Hume’s October 2025 announcement said Octave 2 was approximately 40% faster and half the price of Octave 1, and that it could generate audio in under 200 milliseconds. Current API documentation gives a more specific figure of approximately 100 milliseconds for Octave 2, excluding network transit. Neither figure should be treated as guaranteed end-to-end time to audible audio.

There is also a documentation qualification worth noting: the original Octave material emphasized acting instructions, while the current feature table marks acting instructions for Octave 2 as “coming soon.” That discrepancy may reflect version or documentation changes. Teams should verify the active model’s supported controls before building around them.

See the current model overview and language and feature FAQ.

Voice cloning and voice conversion

Hume says a voice clone can be created from as little as 15 seconds of audio. Octave 2’s launch material also describes cross-language generation intended to preserve the speaker’s accent.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A short recording can be enough to create an initial clone, but it is not proof of studio-grade identity preservation. Test pronunciation, accent transfer, emotional range, consistency, and long-form stability across every target language.

Voice cloning also requires more than technical access. Obtain permission from the speaker, and consider publicity, impersonation, employment, privacy, and disclosure obligations. A paid account does not automatically establish the right to clone a person’s voice.

Rank #4
AI Voice Recorder, Note Voice Recorder - Transcribe & Summarize, AI Noise Cancellation Technology, Supports 152 Languages, 64GB APP Control Audio Recorder for Lectures, Meetings, Calls
  • Cutting-Edge AI Transcription & Summarization: Leverage GPT-4o’s advanced intelligence in this top-tier AI voice recorder for real-time, highly accurate speech-to-text conversion and contextual summarization. Experience natural language processing that delivers polished, instantly usable transcripts—eliminating manual editing. Ideal for professionals seeking efficient documentation
  • 1-Year Unlimited Premium Suite: Unlock 12 months of free DOWAY premium access with your powerful voice recorder: Enjoy limitless transcription, AI-powered professional templates, and smart note-organization tools. Transform recordings into structured documents for business reports, academic notes, or content creation
  • Global 152Language Comprehension: Seamlessly transcribe and summarize content across 152 languages with this intelligent AI recorder – from major business dialects to regional languages. Break communication barriers in international meetings, research, or travel without compromising accuracy
  • Massive 64GB Storage + Military-Grade Cloud Sync: Store 500+ hours of high-fidelity audio internally (no cards needed) on this feature-packed voice recorder, with automatic backups to encrypted cloud storage. Access files securely worldwide through the DOWAY app—your data remains private yet universally available
  • 35-Hour Marathon Battery: Operate this long-lasting voice recorder continuously for 2,100 minutes (35 hours) on one charge. Capture multi-day conferences, field research, or interviews without battery anxiety. Power-optimized for travelers and high-volume users (Note: studio-grade bluetooth 5.3, works Instantly, no Wi-Fi needed)

Octave 2 also supports voice conversion. Hume documents supported input formats including MP3, WAV, M4A, and OGG. Voice conversion is useful when a performance needs to retain timing or delivery while changing vocal identity, but it should be evaluated separately from text-to-speech voice design.

Timestamps for captions and animation

Octave 2 supports word-level and phoneme-level timestamps. These can be used for:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Real-time captions and word highlighting.
  • Avatar lip-sync.
  • Audio segmentation.
  • Post-production editing.
  • Synchronizing speech with animation or other media.

Timestamps must be explicitly requested, and Hume says the feature requires the appropriate Octave 2 request version.

Read the timestamp documentation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to try Octave

No-code testing

Use Hume’s Octave product page or platform playground to try voices from the library, enter text, experiment with voice descriptions, and adjust speed where available. The product page advertises voice selection, cloning, voice design, streaming, multiple audio formats, and timestamp support. Feature availability can depend on the active model and account, so check the current interface rather than assuming every plan exposes every control.

API integration

To use the API, create a Hume account, obtain an API key, store it securely, select a model version, and choose or create a compatible voice. A minimal streaming JSON request based on Hume’s documented pattern is:

curl https://api.hume.ai/v0/tts/stream/json 
  -H "X-Hume-Api-Key: $HUME_API_KEY" 
  -H "Content-Type: application/json" 
  --json '{
    "version": "2",
    "utterances": [
      {
        "text": "I cannot believe you made it.",
        "description": "Deliver this with surprised delight, then soften at the end.",
        "speed": 1.0,
        "trailing_silence": 0.2
      }
    ]
  }'

For a fixed voice, add a voice object to the first utterance:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
"voice": {
  "id": "VOICE_ID"
}

Hume’s voice guide says a voice specified in the first utterance is used for subsequent utterances unless overridden. It also says Octave 1 voices can be used with Octave 1 and Octave 2 requests, while Octave 2 voices require Octave 2.

Best Value
Plaud NotePin S Wearable AI Voice Recorder, Transcribe & Summarize, Black
  • Plaud Intelligence: Capture conversations in 112 languages and generate accurate transcripts with the Plaud App and Web. Plaud Intelligence uses leading models like GPT-5.5, Claude Sonnet 4.6, and Gemini 3.1 Pro to transform raw audio into structured insights. Choose from over 10,000 professional templates to generate mind maps and to-do lists, turning hours of discussion into immediate clarity
  • Multiple Ways To Wear With Included Accessories: Adapt Plaud NotePin S to any workflow instantly with four included accessories. Wear your device effortlessly as a necklace, wristband, clip, or pin. Plaud NotePin S features a dedicated physical record button for precise, tactile control. Stay professional and keep your intelligence within reach all day
  • Enterprise-grade Privacy: Built to the highest standards with ISO 27001/27701, SOC 2, HIPAA, GDPR, and EN18031 compliance. Every conversation is secure and protected. It is the trusted choice for creative, medical, and business professionals handling sensitive info
  • Multimodal Input & Multidimensional Summaries: Capture audio, type notes, add images, and press/tap to highlight for richer context with multimodal input. Press the record button to mark key moments in real time. Plaud transforms a single conversation into multiple perspectives, providing faster, clearer insights, and unifies these inputs to deliver role-specific summaries that reflect your intent and priorities
  • Lightweight Power and Peace of Mind: Weighing only 0.61 oz, Plaud NotePin S delivers 20 hours of continuous recording and 40 days of standby time. Store up to 64GB of audio locally, ensuring you capture every insight even without an internet connection

Verify the live API reference before deployment because endpoints, fields, model compatibility, and output handling can change.

Open the JSON synthesis reference and voice compatibility documentation.

What can developers build?

Octave is relevant to:

  • Narration, audiobooks, podcasts, and voice-over.
  • Game characters and interactive fiction.
  • Animated characters and avatars.
  • Training and instructional content.
  • Conversational interfaces and voice agents.
  • Caption synchronization and lip-sync.
  • Dubbing or accent-preserving voice transformations.

Keep Hume’s TTS and EVI products distinct. TTS converts text into speech. EVI is Hume’s real-time speech-to-speech interface for conversational systems. Octave can be one component in a voice-agent stack, but it is not itself the complete application architecture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pricing and commercial rights

Hume’s pricing page viewed in August 2026 listed the following TTS allowances:

Plan Monthly price shown Included TTS characters Approximate audio
Free $0 10,000 10 minutes
Starter $3 30,000 30 minutes
Creator $7 promotional first month; $14 listed price 140,000 140 minutes
Pro $70 1,000,000 1,000 minutes
Scale $200 3,300,000 3,300 minutes
Business $500 10,000,000 10,000 minutes
Enterprise Custom Custom Custom

The same page listed paid-tier overage rates of $0.15 per 1,000 characters for Creator, $0.12 for Pro, $0.10 for Scale, and $0.05 for Business. The visible table displayed Octave 1 and Octave 2 selectors but did not clearly establish separate pricing for every model and feature. Check the current account interface and terms before budgeting.

Commercial licensing must be checked separately. Hume’s pricing page includes a commercial-license row, but the available information does not establish that every paid plan includes identical commercial rights. Review the current Terms of Use, plan-specific license language, enterprise agreement, and voice-consent requirements.

Check Hume’s current pricing.

Limitations to consider

  • Octave 2 is a preview. Features, prices, quality, and availability can change.
  • Latency is not end-to-end latency. Hume’s model figures exclude network transit and do not guarantee time to first audible audio.
  • Language support is asymmetric. Octave 2 lists 11 synthesis languages, but current voice-design documentation lists English-only voice design.
  • Expressiveness is not perfect controllability. A voice can sound emotional while missing the requested intensity, pause, pronunciation, or attitude.
  • Long-form consistency needs testing. Watch for voice drift, pacing changes, pronunciation errors, and emotional inconsistency.
  • Benchmarks are vendor-reported. Hume’s launch comparison has not been treated here as an independent benchmark.
  • Voice cloning carries legal and ethical risks. Technical capability does not replace consent or rights clearance.
  • Documentation can change. The current acting-instruction status differs from the original launch messaging.

A practical evaluation checklist

Before committing to Octave, test:

  1. The same neutral, emotional, sarcastic, and character dialogue in Octave 1 and Octave 2.
  2. Names, acronyms, numbers, symbols, and uncommon words.
  3. Long-form continuation across scene boundaries.
  4. A cloned voice in each target language.
  5. First-byte latency separately from complete-file generation time.
  6. Word and phoneme timestamp output if captions or animation are required.
  7. Repeatability across several generations of the same line.
  8. Total cost under the intended plan, including overage.
  9. Commercial rights and consent documentation for every real-person voice.

Who should use Octave?

Octave is a strong candidate when contextual delivery, expressive speech, natural-language voice design, short-sample cloning, streaming, or timestamped output matters more than fully deterministic controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It may be less suitable for teams that require a stable non-preview model, independently verified benchmarks, multilingual voice design, strict repeatability, local deployment, frame-level phoneme control, or clearly documented commercial rights on a particular plan.

Teams should compare it with alternatives such as ElevenLabs, Cartesia, and PlayAI using current pricing, language, licensing, cloning, latency, and API documentation.

The Bottom Line

Bottom line: Octave’s meaningful idea is to make text-to-speech more sensitive to semantic context and performance direction. The February 2025 launch established that approach; Octave 2 is the newer, broader, faster-stated preview. It is worth evaluating for expressive characters, narration, voice agents, and developer APIs—but production teams should verify preview limitations, repeatability, language-specific voice design, pricing, and commercial rights before adopting it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.