Model Terminal

GLM-4.6V

GLM-4.6V is Z.ai's multimodal vision-language model that reads and reasons over images, documents, charts, screenshots, and mixed image-text inputs. It adds native multimodal function calling so visual understanding can directly trigger tool actions within agent workflows. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Z.ai
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed