CLAP-HTSAT-Unfused is an open-source contrastive language-audio pretraining model developed by LAION. It learns joint embeddings between audio clips and their textual descriptions. The model is useful for zero-shot audio classification, audio retrieval, and as a feature extractor for downstream audio understanding tasks.
In the Foundation models & chat space, Clap Htsat Unfused takes a focused approach. It focuses on creating unified representations that connect audio content with natural language descriptions. Clap Htsat Unfused is an open-source project aimed at Audio AI researchers and developers. Clap Htsat Unfused is open source under the Open Source license. Clap Htsat Unfused is available on the web and API.
It is developed by LAION, and it first shipped in 2023. Key capabilities include audio embeddings, contrastive learning, and feature extraction.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do