step-3.7-flash

Operational
Ranked #44 in the capability rankings
Vision
Video
Reasoning
Tools
256K

Standard & Paid Price

Standard Price
MeterStandard Price
Input price$0.2000 / 1M tokens
Completion price$1.1500 / 1M tokens
Cache read price$0.0400 / 1M tokens
Billing details
Vendor list price comparison

Description

StepFun step-3.7-flash — sparse MoE with 198B total and 11B active parameters, native 256K context, and native understanding of images and video alongside text. Reasoning depth is selectable per request through reasoning_effort (low / medium / high), and the model supports reliable multi-step tool calling, JSON Schema structured output, streaming, and prompt caching. Tuned for agent and coding workloads. On /v1/responses, send the whole conversation in the input array on every turn: no conversation state is kept upstream, so requests that carry previous_response_id or conversation are refused.

Capabilities

Tasks
chat, reasoning, vision, file-video-understanding

API endpoints

  • POST /v1/messages (anthropic)
  • POST /v1/chat/completions (openai)
  • POST /v1/responses (openai-response)