Model Terminal

NVIDIA Nemotron Nano 12B v2 VL (free)

NVIDIA Nemotron Nano 12B v2 VL is an open 12-billion-parameter vision-language model built on a hybrid Mamba-Transformer architecture, designed for multimodal reasoning over documents, images, tables, charts, and video. It is optimized for high-throughput, low-latency inference on long-context multimodal workloads including OCR, document intelligence, and agentic pipelines. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
NVIDIA
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed