Model Terminal

CLIP ViT-Large/14

CLIP ViT-Large/14 is a vision-language model from OpenAI that encodes images and text into a shared embedding space, enabling zero-shot image classification, image-text similarity scoring, and cross-modal retrieval without task-specific training. It is the large-patch-14 variant of the CLIP family, widely used as a backbone for downstream vision-language research and applications. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Zero-shot-image-classification
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed