- All models
- Reasoning
- Vision
- Audio
All models
All public models
Step 5 Preview
Flagship model for agentic work
Max context
1M
StepFun’s flagship model for agentic work. It delivers frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. It natively supports text, image, and video input and can use tools and documents to carry multi-step tasks through to completion. API model ID: step-5-preview.
Step 3.7 Flash
Flagship multimodal reasoning
Max context
256K
StepFun’s flagship multimodal reasoning model. Building on Step 3.5 Flash’s high-speed reasoning and tool-calling capabilities, it adds native multimodal input and can understand images and videos without an additional vision MCP or auxiliary model. It supports low, medium, and high reasoning effort.
Step 3.5 Flash
Flagship reasoning
Max context
256K
A flagship reasoning model built for agents. It combines deep reasoning with fast responses and stable, reliable tool calling, and is particularly effective for complex project planning and long-horizon task execution.
Step 3.5 Flash 2603
Optimized for agents
Max context
256K
Optimized from Step 3.5 Flash for high-frequency agent workloads. It preserves flagship reasoning and tool calling while improving token efficiency and inference speed, adds a low-reasoning mode to reduce consumption, and improves compatibility with coding tools and agent frameworks.
Step-1o Turbo Vision
Image and video understanding
Max context
32K
Designed for social content creation, AI assistants, and image and video understanding. A single request accepts up to 60 images with a combined size of up to 20 MB, or an MP4 video smaller than 128 MB.
StepAudio 3 Realtime
Natural conversation with reasoning and action
Interaction
Voice to voice
A full-duplex voice flagship for realtime interaction. It combines realtime audio understanding, natural interruption handling, turn coordination, adaptive reasoning, think-while-speaking behavior, and Voice Agent tool calling for in-vehicle systems, smart devices, customer service, and companion applications.
StepAudio 2.5 Chat
Natural, human-like conversation
Output
Text only
A conversational model designed for natural, human-like interaction with text output. It understands complex meaning and paralinguistic cues such as hesitation and laughter, produces emotionally aware responses, and supports fine-grained persona customization.
StepAudio 3 TTS
Human-level speaking performance
The model closely follows natural human speech in timbre, intonation, rhythm, breathing, and pauses. It can express laughter, hesitation, stutters, repetition, and self-correction while adapting emotion and tone to the meaning of the text.
Step TTS Mini
Expressive text-to-speech
Max input
1,000 chars
Supports Mandarin Chinese, English, Japanese, Cantonese, and Sichuanese with 19 official voices. It also supports high-quality voice cloning in Chinese, English, and Japanese for customer outreach, companion applications, and voice assistants that require natural speech.
StepAudio 2.5 ASR
Next-generation streaming ASR flagship
Model scale
4B MTP
A next-generation streaming speech recognition model based on a 4B MTP architecture. It balances recognition accuracy and response latency, supports Chinese and English with ITN text normalization, and is designed for realtime captions, voice input, and meeting transcription.
StepAudio 2 ASR Pro
32B ASR Pro
Parameters
32B
A 32B-parameter professional automatic speech recognition model.
Step ASR
Realtime and offline recognition
Max file size
100 MB
An ASR model with strong Chinese and English recognition. It distinguishes speech from noise and supports mixed Chinese-English speech and heavily accented Mandarin for voice input, voice control, and meeting transcription.
Step-1o Audio
Stable realtime voice
Max interaction
30 min
Provides ultra-low-latency, bidirectional voice conversation. It accepts Chinese, English, and heavily accented Mandarin, and can respond in Chinese, English, Japanese, Cantonese, and Sichuanese. A single interaction can last up to 30 minutes and process up to 70 minutes of audio.

