Llama 4 Maverick
Meta · meta-llama/llama-4-maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Llama 4 Maverick ranks #136 of 348 models in the default AI Benchmark Hub order (quality first, then cost). At $0.45 blended per million tokens, it is cheaper than 71% of models with published API pricing. Its 128,000-token context window is 2.0× smaller than the catalog median (262,144 tokens).
Context
Max output: 115200
Pricing
Output / 1M: 0.70
Blend / 1M: 0.45
Quality
Provider
Moderated: no
Monthly cost at this blended rate
Estimated API spend if every token is billed at $0.45per million (the blend of this model's input and output prices). Real bills depend on the input/output mix.
| Tokens / month | Est. cost (USD) |
|---|---|
| 1,000,000 | $0.45 |
| 10,000,000 | $4.48 |
| 100,000,000 | $44.80 |
Related models
Same lab, similar price, or similar quality — useful next comparisons.
More from Meta
- Llama 3.1 8B Instruct$0.07 / 1M
- Llama 3.2 1B Instruct$0.11 / 1M
- Muse Spark 1.3 Contributor$0.15 / 1M
- Muse Spark 1.2 Contributor$0.15 / 1M
- Llama Guard 4 12B$0.18 / 1M
Closest API price
- Mistral Small 3.1 24B$0.45 / 1M
- Qwen3 Coder Next$0.46 / 1M
- GLM 4.5 Air$0.49 / 1M
- Cydonia 24B V4.1$0.40 / 1M
- Saba$0.40 / 1M