CLIP ViT Base Patch32 is an open-weight vision-language model that maps images and text into a shared embedding space. Developers can use it for image-text similarity, zero-shot image classification, and multimodal retrieval through local model libraries.
Clip Vit Base Patch32 sits in PulseGate's Other AI category. It focuses on matching and classifying images with natural-language descriptions without training a task-specific model. It is built as an open-source project for machine learning developers. The project is open source (MIT). Clip Vit Base Patch32 is available on the web and the command line, and it can be self-hosted.
It is developed by OpenAI, and it first shipped in 2021. The project is developed in the open on GitHub with 34.2k stars. It operates in a well-populated space: PulseGate tracks 8 similar projects. Key capabilities include image-text embeddings, zero-shot classification, and text embeddings.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do