Pricing Details
Pricing for Multimodal Reasoning Models
For
step-5-preview, the cache-miss input price includes writing new content to the cache. Output tokens include both the model’s reasoning process and final answer. See Prompt caching for details.
Pricing for Reasoning Models
Pricing for Vision Models
Image input is counted and billed as tokens, just like text input. By default, each image uses 169 input tokens; with
detail enabled, the token count is calculated from the image size.
Pricing for Speech Models
Billed by token
Billed by characters or audio duration
Text-to-speech
Audio generation and music generation
Speech recognition
For character-based billing, one Chinese character counts as one character, two English letters count as one character, and two punctuation marks count as one character.
Pricing for Image Generation and Editing
step-2x-large and step-image-edit-2 will be retired on October 10, 2026. The text-to-image, image-to-image, and image editing APIs will stop serving requests on that date. See the image model retirement notice.
Image generation is billed by the number of images generated. The default is one image per request.
step-2x-large has not been free since June 12, 2026 and is billed at $0.02 per image.
Pricing for Value-Added Capabilities
Value-added capabilities are billed by actual usage.

