- TTS and ASR — Chinese, English, Japanese, Korean, French, Spanish
- Realtime — Chinese, English
Note: All languages except Chinese and English are in preview. More languages are on the way.
Choose a model
StepAudio 3 TTS
Generate natural, expressive speech—including laughter, hesitation, and self-correction—with low-latency streaming and fine-grained control over context. Learn more about StepAudio 3 TTS.StepAudio 3 ASR
Context-aware transcription for proper names, specialized terminology, dialects, Chinese–English code-switching, and challenging audio such as whispers and singing. Learn more about StepAudio 3 ASR.StepAudio 3 Realtime
Full-duplex real-time voice: build natural, real-time voice agents that handle interruptions, take turns smoothly, reason while speaking, and call tools. Learn more about StepAudio 3 Realtime.StepAudio 3 Gen
Create speech, singing, sound effects, ambience, and background music in a single clip from a natural-language prompt. Learn more about StepAudio 3 Gen.StepAudio 3 Music
Turn lyrics, style prompts, or dry vocals into complete songs, instrumentals, covers, and soundtracks. Learn more about StepAudio 3 Music.StepAudio 2.5 Chat
An end-to-end speech-understanding model served through an OpenAI-compatible Chat Completion API. It accepts audio or text input and returns a text response, interpreting not just the words but also vocal cues such as intonation, hesitation, and laughter.- Paralinguistic understanding: captures a speaker’s emotional state and intent from vocal cues a text transcript would miss.
- Unified text and audio input: typed and spoken turns share the same endpoint, so one conversation can mix both.
- For spoken replies, pass the text output to a text-to-speech model such as StepAudio 2.5 TTS.
StepAudio 2.5 Realtime
An end-to-end speech-to-speech model served over a WebSocket Realtime API. It takes streaming audio or text and responds with streaming audio and a matching text transcript—no separate transcription or synthesis step—keeping latency low enough for natural back-and-forth conversation.- Paralinguistic understanding: interprets vocal cues such as intonation, pacing, and hesitation in addition to the spoken words.
- Natural turn-taking: server-side voice activity detection (VAD) lets the model reply at the right moment and be interrupted mid-response.
- System voices: a set of English system voices selectable per session; see the voices list.
StepAudio 2.5 TTS
A text-to-speech model that brings context understanding into the entire generation pipeline. Instead of matching preset tags, you describe the delivery you want in natural language, at two levels:- Global context sets the overall mood, scene, and character relationships for an entire passage.
- Inline context fine-tunes how individual words and phrases are delivered.
StepAudio 2.5 ASR
A new-generation speech recognition model built on a 4B-parameter Multi-Token Prediction (MTP) architecture, maintaining SOTA transcription accuracy while sharply reducing latency. Supports Chinese and English recognition with ITN text normalization. It offers two access methods:- One-shot recognition (
stepaudio-2.5-asr): submit audio over HTTP + SSE and receive the transcription streamed back incrementally. - Real-time bidirectional streaming (
stepaudio-2.5-asr-stream): stream audio over WebSocket and receive incremental and final results with per-word timing; supports server-side VAD. - Suitable for live captions, voice input, meeting transcription, and back-end batch processing.
Usage limits
- Text-to-speech input: up to 1,000 characters per request.
- Text-to-speech output formats:
wav,mp3,flac,opus,pcm; the default ismp3. - StepAudio 3 Gen input limits: the
instructionandrolesfields are capped at 500 characters andscriptsat 1,000 characters; exceeding a limit returns HTTP 400. - StepAudio 3 Gen delivery: currently supports HTTP
audio/ssereturn; bidirectional WebSocket streaming is not supported.
Next steps
StepAudio 3 TTS
Human-level, low-latency text-to-speech.
StepAudio 3 ASR
Context-aware recognition on a large language model foundation.
StepAudio 3 Realtime
Full-duplex real-time voice with adaptive reasoning and Voice Agent.
StepAudio 3 Gen
Unified audio generation—voice, SFX, ambience, and music.
StepAudio 3 Music
Generate songs, instrumentals, covers, and scoring from dry vocals.
StepAudio 2.5 Chat Overview
Explore the end-to-end speech-understanding model with audio and text input.
StepAudio 2.5 Realtime Overview
Explore the end-to-end speech-to-speech model with real-time conversation over WebSocket.
StepAudio 2.5 TTS Overview
Explore the contextual TTS model with dual-level context control and zero-shot voice cloning.
StepAudio 2.5 ASR Overview
Explore the new-generation 4B MTP ASR model with one-shot SSE and real-time bidirectional streaming.
Voice interaction developer guide
Get started with speech generation, voice cloning, and automatic speech recognition.