Model Terminal

Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's open-weight frontier reasoning model — a 550B total / 55B active parameter Mixture-of-Experts LLM built for agentic workflows including planning, tool use, coding, and long-horizon research tasks. It uses a hybrid Mamba-Transformer MoE architecture with LatentMoE and MTP layers for high-throughput inference on long contexts. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
NVIDIA
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed