
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model supports a native 262K context length and does not implement "thinking mode" (<think> blocks).
Compared to its base variant, this version delivers significant gains in knowledge coverage, long-context reasoning, coding benchmarks, and alignment with open-ended tasks. It is particularly strong on multilingual understanding, math reasoning (e.g., AIME, HMMT), and alignment evaluations like Arena-Hard and WritingBench.
75% off | $0.35$0.0875 | $1.40$0.35 | $0.07$0.0175 | 1.20s | 32 tps | |
|---|---|---|---|---|---|---|
| $0.09 | $0.55 | -- | 0.71s | 11 tps | ||
| $0.09 | $0.58 | -- | 0.76s | 19 tps | ||
| $0.14 | $0.80 | $0.05 | 0.44s | 31 tps | ||
40% off | $0.35$0.21 | $1.40$0.84 | -- | 0.47s | 28 tps | |
| $0.22 | $0.88 | -- | 0.91s | 21 tps | ||
| $0.25 | $1.00 | -- | 0.69s | 20 tps | ||
| $0.1495 | $0.598 | -- | 0.39s | 38 tps | ||
| $0.15 | $0.75 | -- | 0.86s | 9 tps | ||
| $0.20 | $0.60 | -- | 0.85s | 16 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.