Fish Audio
Products
Products
Product overview
Explore the all-in-one creative suite
Voice library
Voices for any role or character
Create
Text to Speech
Generate human-like AI speech
Multi-speaker Dialogue
Create dialogue audio from multi-character scripts
Speech to Text
Transcribe audio and video
Voice Design
Generate custom voices
Voice Changer
Output audio in any voice
Voice Isolation
Extract clear speech
Voice Cloning
Clone your voice
Sound Effects
Coming soon
Generate any sound
Dubbing
Coming soon
Localize audio content
Music
Coming soon
Turn ideas into songs
Images
Generate images from text
Video
Generate video from text or images
S2
Fish Audio S2.1 Pro
Multi-speaker, multi-turn generation with natural language control over voice performance.
API
Platform
OverviewDocsAPI referenceAPI keysAPI pricingAPI Playground
API
Text to Speech
Generate speech through the API
Music
Coming soon
Create songs through the API
Speech to Text
Batch transcribe speech
Sound Effects
Coming soon
Generate sound effects through the API
Real-time Speech to Text
Transcribe speech in real time
Voice Cloning
Clone voices for TTS
Speech Engine
Coming soon
Give agents voice capabilities
Agents
Coming soon
Deploy voice agents in minutes
Dubbing
Coming soon
Translate video and audio through the API
API
Fish Audio API quick start
Debug voice generation, transcription, and account keys online
Text to Speech
Convert text to natural speech with Fish Audio, MiniMax, Qwen, and more
Speech to Text
High-accuracy transcription from uploaded audio
Voice Cloning
Clone your voice in about a minute from short samples
Voice Gallery
Browse public models and pick a reference voice
AI Image
Generate images from prompts with leading models
AI Video
Create video from text descriptions and styles
Lip-sync & digital human
Align speech to video for avatars and presenters
Voice Workspace
Voice synthesis workspace to create and manage your voice projects
Short video & dubbing
Fast voiceover for social, ads, and UGC
Audiobooks & podcasts
Long-form narration with natural pacing
Education & training
Clear narration for courses and internal comms
Company
AboutBlog
Resources
Coze
Tavo
SillyTavern
Dify
Open WebUI
AnythingLLM
Home Assistant
n8n
Affiliate program
Coming soon
API Playground
Try REST endpoints online with your API key
API keys
Create and manage API keys in your account
Pricing
TTS Model Guide2026-03-19·6 min read

Xiaomi MiMo-V2-TTS: Text-Driven Expressive TTS

From free-form style instructions to non-verbal events and singing capability: MiMo-V2-TTS brings “expression” into speech generation.

Read Official Details →Open API Playground →

Why MiMo Is More Than Traditional TTS

Free-Form Style Instructions

Describe emotion, pacing, tone, and performance intent in natural language; the model parses it into generation behavior.

Contextual Emotion & Prosody

Not just labels—MiMo adapts intonation and rhythm based on text semantics and context.

Natural Non-Verbal Events

Pauses, hesitation fillers, sighs, coughing, and laughter are integrated into the generation process.

Singing in One Unified Model

The official page highlights singing capability within the same unified model.

If your site is built around voice-over workflows, MiMo’s value is that you can encode performance and emotional details directly into text—without relying on rigid UI dropdowns.

How to Write Style Prompts (Reusable in Your Workflow)

A practical template: emotion/pacing/performance intensity + voice tone/texture + (optional) non-verbal events.

Quick Example

angry but trying to stay calm, slightly clipped delivery, quick pacing

Whisper / Soft

deeply affectionate, speaking slowly, almost whispering, warm and soft

Add Non-Verbal Events

Hold on... [heavy breathing] I... I need a minute... [soft cough] Just give me time.

Treat these prompts as your site’s “style templates”. When you integrate MiMo later, you just map the templates to MiMo’s fields/control interface.

Product Integration: Suggested Field Mapping

1) Add MiMo as a provider/model

Add MiMo in your model configuration (e.g., `provider=mimo` and a model identifier) so users can select it in the model picker.

2) Normalize “free style prompts” into your input

If you already have `emotion/language/speed/volume`, use a “prompt composition” strategy to build MiMo’s style description; or add a dedicated `stylePrompt` field for direct pass-through.

3) Handle billing/quota based on text size

Providers have different generation costs. Start with a character-based or estimated-duration multiplier, then calibrate using real generation metrics.

4) Reduce learning cost with docs & examples

After SEO brings users in, the most important thing is “copyable prompts”. Provide examples, marker explanations, and common Q&A.

FAQ

How does MiMo-V2-TTS control style?

It emphasizes free-form style instructions: you describe emotion, pacing, tone, and performance intent in natural language instead of selecting a fixed emotion tag.

Can it generate pauses, breathing, coughing, and laughter?

The official page shows text markers that guide non-verbal events like pauses, hesitation fillers, sighs, coughing, and laughter for more natural performance.

Does MiMo-V2-TTS support singing?

Yes. The official description highlights native singing capability within the same unified model.

When will FishSpeech support MiMo-V2-TTS?

We are planning provider integration, field mapping, and billing/quotas. For now, use this article to build your prompt style spec, and keep an eye on upcoming documentation updates.

Related

Style prompt templates →Non-verbal events →Singing capability →API Playground →Text-to-Speech Tool →Voice Clone Tutorial →

Product

  • Text to Speech
  • Multi-speaker Dialogue
  • Speech to Text
  • Voice Design
  • Voice Changer
  • Voice Isolation
  • Voice Cloning
  • Sound EffectsComing soon
  • DubbingComing soon
  • MusicComing soon
  • Images
  • Video

Solutions

  • Short video & dubbing
  • Audiobooks & podcasts
  • Education & training

Research

  • Fish Audio S1
  • Fish Audio S2 Pro
  • Fish Audio S2.1 Pro

Resources

  • Docs
  • API reference
  • Model library
  • Voice clone tutorial
  • Product comparison

Company

  • About
  • Blog
Kitta AudioPowered by Fish Audio
© 2026 Kitta AI. All rights reserved.|Privacy Policy|Terms of Service|Report abuse|support@fishaudio.org
support@fishaudio.org

Kitta Audio is independently operated by Kitta AI, not the official Fish Audio service.