Model information
Capabilities
- Human-level speaking performance: in timbre, intonation, rhythm, and breath, the model follows natural human speech—laughter, lip smacks, hesitation, stutters, repetition, self-correction—and shifts emotion and tone to fit the context, sounding like a person rather than a synthesized clip.
- Ultra-low-latency streaming: it outputs while generating, so playback can start before a full sentence is synthesized—ideal for real-time conversation and voice assistants.
API endpoints
Non-streaming synthesis
POST /v1/audio/speechStreaming synthesis
WebSocket /v1/realtime/audioRelated resources
Audio models overview
Back to the Audio 3 model overview.
StepAudio 3 Gen
Unified audio generation—voice, SFX, ambience, and music.
Full pricing details
Billing rules for all speech, text, and image models.
Voice list
Official voices and parameter notes.
Voice Studio
Try the full capabilities of StepAudio 3 TTS online.