This is an OpenCLIP model based on ConvNeXt architecture trained on the LAION-400M dataset with 13B samples seen. It produces embeddings for images and text that can be used for zero-shot classification, image-text retrieval, and other multimodal tasks. The model is provided as open weights on Hugging Face for use with the OpenCLIP library and is intended for researchers and developers building vision-language applications.
CLIP Convnext Base laion400M s13B b51K sits in PulseGate's Multimodal & vision category. It focuses on training and deploying open-source vision-language models without proprietary data or closed weights. It is built as an open-source project for machine learning researchers and developers. The project is open source (Open Source). It runs on the web and API.
LAION builds and maintains CLIP Convnext Base laion400M s13B b51K. Key capabilities include vision-language embedding, zero-shot classification, and image-text retrieval.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do