Model Terminal

CLIP ViT-B/32

CLIP ViT-B/32 is OpenAI's vision-language embedding model that encodes images and text into a shared vector space for zero-shot image classification, image-text retrieval, and multimodal similarity tasks. It uses a Vision Transformer (ViT-B/32) image encoder paired with a text transformer encoder, trained via contrastive learning on 400 million image-caption pairs. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Zero-shot-image-classification
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed