
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings.
Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating.
| $0.13 | $0.53 | $0.033 | 1.42s | 35 tps | ||
| $0.132 | $0.528 | $0.033 | 3.66s | 57 tps | ||
| $0.14 | $0.58 | $0.035 | 4.40s | 46 tps | ||
| $0.14 | $0.58 | $0.035 | 2.58s | 50 tps | ||
| $0.15 | $0.64 | $0.04 | 2.49s | 66 tps | ||
| $0.20 | $0.80 | $0.05 | 2.73s | 172 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.