Fish Audio
Products
Products
Product overview
Explore the all-in-one creative suite
Voice library
Voices for any role or character
Create
Text to Speech
Generate human-like AI speech
Multi-speaker Dialogue
Create dialogue audio from multi-character scripts
Speech to Text
Transcribe audio and video
Voice Design
Generate custom voices
Voice Changer
Output audio in any voice
Voice Isolation
Extract clear speech
Voice Cloning
Clone your voice
Sound Effects
Coming soon
Generate any sound
Dubbing
Coming soon
Localize audio content
Music
Coming soon
Turn ideas into songs
Images
Generate images from text
Video
Generate video from text or images
S2
Fish Audio S2.1 Pro
Multi-speaker, multi-turn generation with natural language control over voice performance.
API
Platform
OverviewDocsAPI referenceAPI keysAPI pricingAPI Playground
API
Text to Speech
Generate speech through the API
Music
Coming soon
Create songs through the API
Speech to Text
Batch transcribe speech
Sound Effects
Coming soon
Generate sound effects through the API
Real-time Speech to Text
Transcribe speech in real time
Voice Cloning
Clone voices for TTS
Speech Engine
Coming soon
Give agents voice capabilities
Agents
Coming soon
Deploy voice agents in minutes
Dubbing
Coming soon
Translate video and audio through the API
API
Fish Audio API quick start
Debug voice generation, transcription, and account keys online
Text to Speech
Convert text to natural speech with Fish Audio, MiniMax, Qwen, and more
Speech to Text
High-accuracy transcription from uploaded audio
Voice Cloning
Clone your voice in about a minute from short samples
Voice Gallery
Browse public models and pick a reference voice
AI Image
Generate images from prompts with leading models
AI Video
Create video from text descriptions and styles
Lip-sync & digital human
Align speech to video for avatars and presenters
Voice Workspace
Voice synthesis workspace to create and manage your voice projects
Short video & dubbing
Fast voiceover for social, ads, and UGC
Audiobooks & podcasts
Long-form narration with natural pacing
Education & training
Clear narration for courses and internal comms
Company
AboutBlog
Resources
Coze
Tavo
SillyTavern
Dify
Open WebUI
AnythingLLM
Home Assistant
n8n
Affiliate program
Coming soon
API Playground
Try REST endpoints online with your API key
API keys
Create and manage API keys in your account
Pricing
Voice Clone Tutorial

Text to Speech Fine-grained Control

Advanced control over speech generation

Getting Started

Disabling normalization may reduce the stability of reading numbers, dates, and URLs. You'll need to handle these cases manually for best results.

Phoneme Control

Phoneme control allows you to specify exact pronunciations for words or characters. Currently, we support:

  • CMU Arpabet (for English)
  • Pinyin (for Chinese)

To use phoneme control, wrap the desired pronunciation in <|phoneme_start|> and <|phoneme_end|> tags. Each tag should contain a single word or character.

Examples

Standard: I am an engineer.

With control: I am an <|phoneme_start|>EH N JH AH N IH R<|phoneme_end|>.

标准: 我是一个工程师。

控制: 我是一个<|phoneme_start|>gong1<|phoneme_end|><|phoneme_start|>cheng2<|phoneme_end|><|phoneme_start|>shi1<|phoneme_end|>。

Paralanguage

Paralanguage controls allow you to add natural speech elements and pauses to make the generated speech sound more human-like. There are two main types of controls:

Pause Words

You can use common pause words like "um", "uh", "嗯", "啊" to control the rhythm of the speech.

Special Effects

The following special effects can be added using parentheses:

EffectDescriptionFirst AvailableStage
(break)Short pauseV2Experimental
(long-break)Extended pauseV2Experimental
(breath)Breathing soundV2Experimental
(laugh)Laughter soundV2Experimental
(cough)Coughing soundV2Experimental
(lip-smacking)Lip smacking soundV2Experimental
(sigh)Sighing soundV2Experimental

The effects (laugh), (cough), (lip-smacking), and (sigh) are developing. You may need to repeat them multiple times for better results.

English Example:

Standard: I am an engineer.

With paralanguage: I am, um, an (break) engineer.

中文示例:

标准: 我是一名工程师。

添加副语言: 我,嗯,是一名(break)工程师。

Next Steps

API Playground →
Integrate via API
Use the text-to-speech API to build voice-powered apps. Try it live in the playground.
Try Free →
Voice Cloning Free Online
Create an authorized voice model in under 1 minute — no credit card required. Start with the free tier today.