
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts).
Pricing for the Alibaba endpoints varies by context length. Once a request is greater than 128k input tokens, the higher pricing is used.
| $0.22 | $1.80 | -- | 0.94s | 67 tps | ||
| $0.30 | $1.00 | $0.10 | 0.51s | 36 tps | ||
| $0.38 | $1.55 | -- | 1.58s | 8 tps | ||
| $0.975 | $4.875 | -- | 1.57s | 29 tps | ||
| $0.35 | $1.50 | $0.04 | 0.91s | 21 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
