Quick links
Model capability overview
Browse public models across five capability categories
- Recommended models
- All models
- Reasoning
- Vision
- Audio
Recommended
Recommended models
Step 5 Preview
Flagship model for agentic work
Max context
1M
StepFun’s flagship model for agentic work. It delivers frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. It natively supports text, image, and video input and can use tools and documents to carry multi-step tasks through to completion. API model ID: step-5-preview.
Step 3.7 Flash
Flagship multimodal reasoning
Max context
256K
StepFun’s flagship multimodal reasoning model. Building on Step 3.5 Flash’s high-speed reasoning and tool-calling capabilities, it adds native multimodal input and can understand images and videos without an additional vision MCP or auxiliary model. It supports low, medium, and high reasoning effort.
Step 3.5 Flash 2603
Optimized for agents
Max context
256K
Optimized from Step 3.5 Flash for high-frequency agent workloads. It preserves flagship reasoning and tool calling while improving token efficiency and inference speed, adds a low-reasoning mode to reduce consumption, and improves compatibility with coding tools and agent frameworks.
StepAudio 3 Realtime
Natural conversation with reasoning and action
Interaction
Voice to voice
A full-duplex voice flagship for realtime interaction. It combines realtime audio understanding, natural interruption handling, turn coordination, adaptive reasoning, think-while-speaking behavior, and Voice Agent tool calling for in-vehicle systems, smart devices, customer service, and companion applications.
StepAudio 2.5 Chat
Natural, human-like conversation
Output
Text only
A conversational model designed for natural, human-like interaction with text output. It understands complex meaning and paralinguistic cues such as hesitation and laughter, produces emotionally aware responses, and supports fine-grained persona customization.
StepAudio 3 TTS
Human-level speaking performance
The model closely follows natural human speech in timbre, intonation, rhythm, breathing, and pauses. It can express laughter, hesitation, stutters, repetition, and self-correction while adapting emotion and tone to the meaning of the text.
StepAudio 2.5 ASR
Next-generation streaming ASR flagship
Model scale
4B MTP
StepFun’s next-generation speech recognition model uses a 4B-parameter Multi-Token Prediction architecture to predict multiple tokens in parallel. It transcribes five minutes of audio in under one second while maintaining state-of-the-art accuracy, making it suitable for Voice Agents, batch transcription, realtime captions, and live streaming.
Step 3.5 Flash
Flagship reasoning
Max context
256K
A flagship reasoning model built for agents. It combines deep reasoning with fast responses and stable, reliable tool calling, and is particularly effective for complex project planning and long-horizon task execution.
Step TTS Mini
Expressive text-to-speech
Max input
1,000 chars
Supports Mandarin Chinese, English, Japanese, Cantonese, and Sichuanese with 19 official voices. It also supports high-quality voice cloning in Chinese, English, and Japanese for customer outreach, companion applications, and voice assistants that require natural speech.
Step Image Edit 2
Unified text-to-image and image editing
Edit latency
1-2s
StepFun’s latest lightweight editing model supports both text-to-image generation and image editing in a single model. With fewer than 6B parameters, it delivers strong performance and enables responsive, interactive image editing.

