Skip to main content
step-3.5-flash is StepFun’s flagship language reasoning model. It delivers top-tier reasoning quality and fast, reliable execution — decomposing and planning complex tasks, and reliably orchestrating tool calls. Suitable for logical reasoning, math, software engineering, deep research, and other complex workloads. Built on a 196B-parameter / 11B-activation sparse MoE architecture.

Key facts

Model type

Sparse MoE architecture
196B total params / 11B activated params

Context length

256K tokens

Best for

High-throughput reasoning + tool calling
Optimized for agent and coding workloads

Core capabilities

🚀 High-throughput reasoning

Sparse MoE architecture delivers high throughput and low latency, ideal for real-time agent workflows and high-volume calls.

🛠️ Tool calling

Reliable tools / tool_choice orchestration, supports multi-step task decomposition and plan execution.

🧠 Complex reasoning

Handles logical reasoning, math, software engineering, and deep research — a dependable foundation for long-chain agent reasoning.

Model variants

step-3.5-flash

Base version: general-purpose reasoning and tool calling — well-suited to most Agent and complex-task scenarios.

step-3.5-flash-2603

Agent-optimized version: tuned from step-3.5-flash for high-frequency agent scenarios. Better token efficiency and faster inference, with an optional low-reasoning mode that significantly reduces token consumption. Improved compatibility with coding workflows and agent frameworks. Supports the reasoning_effort field (low / high).

API endpoint

Chat Completion

POST /v1/chat/completions
OpenAI-compatible, with streaming and tool calling.

Pricing

See full pricing details →

Quickstart

Reasoning model guide

Recommended usage of reasoning models for complex tasks, tool calling, and long contexts.

Step 3.7 Flash — multimodal flagship

Building on Step 3.5 Flash with native image and video understanding.