Model Terminal

UniDepthV2 ViT-L/14

UniDepthV2 ViT-L/14 is a large-scale monocular metric depth estimation model that predicts real-world depth from a single RGB image without requiring camera intrinsics at inference time. It is the highest-capacity checkpoint in the UniDepthV2 family, offering metric 3D scene estimates with per-pixel confidence outputs across diverse camera types and scenes. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Text
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed