Skip to main content
Clone a voice from a previously uploaded WAV or MP3 file so it can be used for TTS audio generation or Realtime voice conversation.
Cloned voices can be used only with language omitted, zh, or en. Calling Text-to-Speech with any other language and a cloned voice returns 400 input_invalid (voice <voice_id> does not support language <language>). For other languages, use a system voice that supports the language (see Get system voices).

Endpoint

POST https://api.stepfun.ai/v1/audio/voices

Request parameters

  • model string required
    Cloning model. Use stepaudio-2.5-tts to clone a voice for use with StepAudio 2.5 TTS or Realtime.
  • text string optional
    Transcript of the source audio file. If omitted, automatic speech recognition is used. For best results, we recommend providing the transcript.
  • file_id string required
    File ID of the source audio used for cloning. Obtain the ID via file upload; set purpose to storage. Supported formats: mp3, wav. Audio length should be 5–10 seconds.
  • sample_text string optional
    Text (max 50 characters) used to create a preview clip.

Response

  • id string
    Voice ID for subsequent audio generation.
  • object string
    Object type, always audio.voice.
  • duplicated boolean
    Indicates the request was duplicated (returned on repeated calls).
  • sample_text string
    Text used for the preview audio.
  • sample_audio string
    Preview audio in base64 (wav). Convert to a file to play.

Example