Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token. The model supports a 256K context window and exposes selectable reasoning levels (high/medium/low), letting callers trade off speed, cost, and depth of reasoning.
Designed for coding, agentic workflows, structured outputs, and long-context productivity tasks.
| $0.20$0.16 | $1.15$0.92 | $0.04$0.032 | 0.45s | 98 tps | ||
| $0.20 | $1.15 | $0.04 | 2.07s | 8 tps | ||
| $0.20 | $1.15 | $0.04 | 3.42s | 91 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.