Model Terminal

Moondream2

Moondream2 is a compact open-source vision-language model designed for efficient on-device and edge inference. It accepts image and text inputs to perform visual question answering, image captioning, object detection, and region pointing. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed