Recommended
4
Public models
5
Max context
256K
4 category views covering 5 public models
- Recommended models
- All models
- Text & reasoning
- Audio
Recommended
Recommended models
Step 3.7 Flash
Flagship multimodal reasoning
Max context
256K
StepFun’s flagship multimodal reasoning model. Building on step-3.5-flash’s high-throughput reasoning and tool calling, it adds native multimodal input: understanding images and videos directly, without an additional vision MCP or auxiliary model. Three reasoning effort levels (low / medium / high) make it a fast and dependable choice for agent, coding, and multimodal workloads.
step-3.5-flash
Flagship reasoning
Max context
256K
A flagship reasoning model built for agents, combining deep reasoning with ultra-fast responses and stable, reliable tool calling. On top of strong general reasoning, it excels at complex project planning and long-horizon task execution.
stepaudio-3-tts
Human-level speaking performance
Max input
1,000 chars
In timbre, intonation, rhythm, and breath, the model tracks natural human speech—laughter, hesitation, stutters, repetition, self-correction—and shifts emotion and tone to match the context. Ideal for real-time conversation, voice assistants, and high-expressiveness content.
stepaudio-2.5-chat
End-to-end speech understanding
Input
Audio + text
An end-to-end speech-understanding model served through an OpenAI-compatible Chat Completion API. It accepts audio or text input and returns text, interpreting vocal cues such as intonation, hesitation, and laughter alongside the words. It suits voice assistants, call and meeting analysis, and voice message triage.
Related entry points