Step 3.7 Flash
Stepfun
stepfun/step-3.7-flashStep 3.7 Flash is StepFun's multimodal vision-language model, built on Step 3.5 Flash with an added vision encoder for native image understanding. With a 256k context window, selectable reasoning levels, and high throughput, it targets agentic, coding, and search workflows that mix text and visual input.
Cost rate
0.2x
Context
262K
Released
May 28, 2026
Input
TextImageVideo
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
No top-ranked categories found.
Performance
Median latency and throughput measured across recent requests.
Throughput117 tok/s
Latency