This is a larger variant of the CLAP (Contrastive Language-Audio Pretraining) model from LAION, specialized for music and speech. It maps audio and text into a shared embedding space, enabling zero-shot classification, retrieval, and other audio understanding applications. The model is available on Hugging Face.
In the Foundation models & chat space, Larger Clap Music And Speech takes a focused approach. It focuses on creating unified embeddings for music and speech audio for retrieval and classification tasks. It is built as an open-source project for audio AI researchers and developers. The project is open source (Open Source). It runs on the web and API.
It is developed by LAION, and it first shipped in 2023. Key capabilities include Audio Feature Extraction, Music Understanding, and Speech Understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do