Model Terminal
Qwen2.5-VL-3B-Instruct
Qwen2.5-VL-3B-Instruct is a 3-billion-parameter instruction-tuned vision-language model from Alibaba Cloud's Qwen team, capable of understanding images, documents, charts, and video. It is the smallest variant in the Qwen2.5-VL family, designed for efficient deployment across multimodal tasks including document analysis, visual question answering, structured output generation, and agentic UI interaction. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Hugging Face
- Openness
- —
- Modalities
- Image-text-to-text
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —