Model Terminal

CLAP HTSAT-Fused

CLAP HTSAT-Fused is an open-source contrastive audio-language pretraining model from LAION that maps audio clips and text prompts into a shared embedding space. It supports zero-shot audio classification, audio-text retrieval, and feature extraction, using feature fusion to handle variable-length audio inputs. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Audio-classification
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed