Model Terminal

Qwen2-VL-7B-Instruct

Qwen2-VL-7B-Instruct is a 7-billion-parameter instruction-tuned vision-language model from Alibaba's Qwen team, capable of understanding images, long-form video, and multilingual text within visual content. It supports tasks such as document understanding, visual question answering, and agentic visual reasoning. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Qwen
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed