Skip to main content
StepAudio 3 Gen creates comprehensive audio content from natural language. Describe roles, timbre, and emotion, and it generates expressive, controllable voice—then adds sound effects, ambient sound, background music, and singing, orchestrated into a single finished clip. Technical report · More demos

Model information

stepaudio-3-gen-preview is the model name used during the free-trial period. When the free trial ends, this preview version will be retired and a formal paid version added.

Capabilities

  • Fine-grained voice design and control: control multi-role lines, timbre, speaking manner, emotion, and dialect through natural language, along with paralinguistic features like laughter, breathing, and pauses. Describe several roles and their interactions in one instruction, and the model organizes the dialogue and transitions for coherent, expressive results.
  • All-element generation and temporal orchestration: beyond voice, it also generates sound effects, ambient sound, background music, and singing, and lets you specify where and in what order each element appears—coordinated generation and seamless clip stitching for complete audio productions.

Use cases

  • Professional content production: for film, games, animation, radio drama, and audiobooks—design each role’s timbre and delivery to quickly produce character dubbing and multi-role dialogue, plus the ambient sound, sound effects, and background music you need, cutting the coordination and post-production cost of traditional workflows.
  • Short video and self-media: for creators’ short-video dubbing—no professional software or audio experience required; a single natural-language description generates style-controllable voiceover that matches the video, with background music and sound effects added on demand.

API endpoint

Audio Generation API

POST /v1/audio/generate

Audio models overview

Back to the Audio 3 model overview.

StepAudio 3 TTS

Human-level text-to-speech from the same series.

Full pricing details

Billing rules for all speech, text, and image models.

Voice Studio

Try StepAudio 3 Gen online.