Skip to main content
From a natural-language description, lyrics, vocals, or reference audio, the model covers the full arc of music creation—from a rough idea to a finished song to intelligent scoring. View more demos

Model information

stepaudio-3-music-preview is the model name used during the free-trial period. When the free trial ends, this preview version will be retired and a formal paid version added.

Model architecture

StepMusic uses a unified architecture for music understanding and high-fidelity generation. Trained on large-scale song and music-understanding data, it learns melody, harmony, rhythm, structure, timbre, and vocal information through a self-developed unsupervised music-representation encoder, and completes the final audio synthesis with generation and rendering modules.
  • Music semantic understanding: maps text descriptions, lyrics, and reference audio into a unified music semantic space, understanding the relationships among content, style, and vocal requirements.
  • Multi-condition control: combines conditions such as lyrics, music description, vocals, melody, or a reference song by task, supporting multiple generation modes.
  • Music structure modeling: learns structures such as intro, verse, chorus, bridge, interlude, and outro, along with melodic and energy development across sections.
  • High-fidelity audio rendering: reconstructs audio detail through generation architectures such as DiT, balancing vocals, instruments, spatiality, and overall production quality.

Core capabilities

  • Lyrics to song: from a music description and lyrics, generate a complete song with vocals, melody, arrangement, and mixing.
  • Instrumental generation: from a style, mood, scene, or instrument description, generate instrumental music with no vocals.
  • Intelligent scoring from dry vocals: analyze the melody, rhythm, and mood of an a cappella or dry vocal and automatically generate a matching accompaniment and arrangement.
  • Natural-language music control: describe style and genre, vocals and singing style, melody and rhythm, instruments and arrangement, mood and scene, and production quality.
  • Reference-audio driven: use timbre, humming, an original track, or dry vocals as reference input for more precise, more personalized generation.
  • ABC notation input: music generation from ABC melody notation (coming soon).

Use cases

  • Music creation and demos: quickly validate lyrics, melody, genre, and arrangement directions, providing draft material for songwriters, producers, and singers.
  • Short video and content platforms: generate original songs or background music by video theme, mood, and duration, improving content-production efficiency.
  • Film, games, and advertising: generate theme music, scene music, and ad music around characters, plot, world-building, or brand tone.
  • Virtual characters and digital humans: generate songs with a character’s timbre to strengthen the character’s vocal identity and sustain content supply.
  • Personalized entertainment: generate custom songs using your own voice, lyrics, or humming for keepsakes, greetings, and social sharing.
  • Music education and inspiration: quickly turn a melodic idea into versions in different styles, helping to understand arrangement, genre, and emotional expression.

Input tips

  • Be specific: rather than “generate a nice song,” also state the genre, tempo, vocals, instruments, mood, and production quality to more easily get results that match your expectations.
  • Use clear lyric structure: use section tags such as [Intro], [Verse], [Pre-Chorus], [Chorus], [Bridge], and [Outro] to help the model understand the song structure.
  • Adjust key variables round by round: settle the overall genre and vocal direction first, then gradually adjust instruments, arrangement density, and production detail to compare versions.

Compliance notes

For voice cloning, song covers, and reference-audio generation, ensure you have obtained lawful authorization from the voice subject and the relevant rights holders of the lyrics, composition, and recording. Before public distribution or commercial use, complete the necessary copyright, personality-rights, and content-safety review according to your region, platform rules, and use case.

API endpoint

Music API

POST /v1/audio/music

Audio models overview

Back to the Audio 3 model overview.

Full pricing details

Billing rules for all speech, text, and image models.

Music API

View the Music API request parameters.

Voice Studio

Try the music generation of StepAudio 3 Music online.