Model Terminal
CLIP ViT-B/32
CLIP ViT-B/32 is OpenAI's vision-language embedding model that encodes images and text into a shared vector space for zero-shot image classification, image-text retrieval, and multimodal similarity tasks. It uses a Vision Transformer (ViT-B/32) image encoder paired with a text transformer encoder, trained via contrastive learning on 400 million image-caption pairs. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Hugging Face
- Openness
- —
- Modalities
- Zero-shot-image-classification
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —