Model Terminal

Qwen2.5-VL-3B-Instruct

Qwen2.5-VL-3B-Instruct is a 3-billion-parameter instruction-tuned vision-language model from Alibaba Cloud's Qwen team, capable of understanding images, documents, charts, and video. It is the smallest variant in the Qwen2.5-VL family, designed for efficient deployment across multimodal tasks including document analysis, visual question answering, structured output generation, and agentic UI interaction. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Image-text-to-text
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed