Nemotron 3 Ultra
NVIDIA · nvidia/nemotron-3-ultra-550b-a55b
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Nemotron 3 Ultra ranks #248 of 348 models in the default AI Benchmark Hub order (quality first, then cost). At $1.88 blended per million tokens, it is cheaper than 36% of models with published API pricing. Its 256,000-token context window is about the catalog median (262,144 tokens).
Context
Max output: 32768
Pricing
Output / 1M: 3.13
Blend / 1M: 1.88
Quality
Provider
Moderated: no
Monthly cost at this blended rate
Estimated API spend if every token is billed at $1.88per million (the blend of this model's input and output prices). Real bills depend on the input/output mix.
| Tokens / month | Est. cost (USD) |
|---|---|
| 1,000,000 | $1.88 |
| 10,000,000 | $18.75 |
| 100,000,000 | $187.50 |
Related models
Same lab, similar price, or similar quality — useful next comparisons.
More from NVIDIA
- Nemotron 3 Nano Omni$0.00 / 1M
- Nemotron 3 Nano 30B A3B$0.12 / 1M
- Nemotron 3.5 Lightning$0.14 / 1M
- Nemotron 3.5 Content Safety$0.20 / 1M
- Nemotron 3 Super$0.24 / 1M
Closest API price
- Grok 4.3$1.88 / 1M
- Grok 4.20 Multi-Agent$1.88 / 1M
- Grok 4.20$1.88 / 1M
- KAT-Coder-Pro V2.5$1.85 / 1M
- Qwen3 Coder Plus$1.95 / 1M