Model information
stepaudio-3-gen-preview is the model name used during the free-trial period. When the free trial ends, this preview version will be retired and a formal paid version added.Capabilities
- Fine-grained voice design and control: control multi-role lines, timbre, speaking manner, emotion, and dialect through natural language, along with paralinguistic features like laughter, breathing, and pauses. Describe several roles and their interactions in one instruction, and the model organizes the dialogue and transitions for coherent, expressive results.
- All-element generation and temporal orchestration: beyond voice, it also generates sound effects, ambient sound, background music, and singing, and lets you specify where and in what order each element appears—coordinated generation and seamless clip stitching for complete audio productions.
Use cases
- Professional content production: for film, games, animation, radio drama, and audiobooks—design each role’s timbre and delivery to quickly produce character dubbing and multi-role dialogue, plus the ambient sound, sound effects, and background music you need, cutting the coordination and post-production cost of traditional workflows.
- Short video and self-media: for creators’ short-video dubbing—no professional software or audio experience required; a single natural-language description generates style-controllable voiceover that matches the video, with background music and sound effects added on demand.
API endpoint
Audio Generation API
POST /v1/audio/generateRelated resources
Audio models overview
Back to the Audio 3 model overview.
StepAudio 3 TTS
Human-level text-to-speech from the same series.
Full pricing details
Billing rules for all speech, text, and image models.
Voice Studio
Try StepAudio 3 Gen online.