
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.
Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
| $0.10 | $0.32 | -- | 0.50s | 14 tps | ||
| $0.20 | $0.52 | $0.10 | 1.33s | 21 tps | ||
| $0.22 | $0.50 | $0.11 | 0.86s | 33 tps | ||
| $0.293 | $2.253 | -- | 0.47s | 29 tps | ||
25% off | $0.60$0.45 | $1.20$0.90 | -- | 0.61s | 28 tps | |
| $0.59 | $0.79 | $0.295 | 0.27s | 143 tps | ||
| $0.71 | $0.71 | $0.71 | 0.94s | 50 tps | ||
| $0.72 | $0.72 | -- | 1.34s | 129 tps | ||
| $0.72 | $0.72 | -- | -- | -- | ||
| $1.04 | $1.04 | -- | 0.77s | 19 tps | ||
| $0.135 | $0.40 | -- | 1.34s | 15 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.