Skip to main content
StepAudio 3 TTS produces natural, human-like speech from text, with fine-grained control over delivery and low-latency streaming. Technical report

Model information

Capabilities

  • Human-level speaking performance: in timbre, intonation, rhythm, and breath, the model follows natural human speech—laughter, lip smacks, hesitation, stutters, repetition, self-correction—and shifts emotion and tone to fit the context, sounding like a person rather than a synthesized clip.
  • Ultra-low-latency streaming: it outputs while generating, so playback can start before a full sentence is synthesized—ideal for real-time conversation and voice assistants.
Suitable for real-time conversation, voice assistants, and high-expressiveness content generation.

API endpoints

Non-streaming synthesis

POST /v1/audio/speech

Streaming synthesis

WebSocket /v1/realtime/audio

Audio models overview

Back to the Audio 3 model overview.

StepAudio 3 Gen

Unified audio generation—voice, SFX, ambience, and music.

Full pricing details

Billing rules for all speech, text, and image models.

Voice list

Official voices and parameter notes.

Voice Studio

Try the full capabilities of StepAudio 3 TTS online.