Fish Audio Drama 3 is coming soon: direct speech in plain language
Fish Audio has shared a preview of Drama 3, a text-to-speech model focused on directing a performance in plain language rather than learning a set of audio tags. The capabilities below come from the upstream preview announcement; they are not available here or independently verified results.
Coming soonThis page covers a model preview. Drama 3 is not connected to this site’s text-to-speech generator or model picker. Availability, pricing, API details and launch timing depend on future Fish Audio announcements.
What the announcement describes
Describe the performance
Use ordinary language to describe tone, pacing and character without first choosing fixed audio tags.
Shift mid-sentence
Change the direction of a voice performance partway through a sentence.
Stage multiple characters
Arrange dialogue for different characters within one scene.
Revise one word
The preview describes correcting a single word instead of redoing the whole passage.
Where this could help
Character dialogue, podcasts and narrated stories could benefit from more direct performance direction and local revisions. Real-world results and supported workflows still need to be evaluated after access opens.
Can I try it here?
This page covers a model preview. Drama 3 is not connected to this site’s text-to-speech generator or model picker. Availability, pricing, API details and launch timing depend on future Fish Audio announcements.