Audio models
What model should I use for text-to-speech?
Usestepaudio-2.5-tts — our flagship contextual TTS model with zero-shot voice cloning and natural-language control over emotion and style through the instruction field and inline () prompts. See Audio Models for details.
Where can I find the TTS API parameters?
See Generate audio for the full request schema and examples.What audio formats are supported?
wav, mp3, flac, opus, pcm. Default is mp3.