chinese-clip-vit-base-patch16 is a Vision Transformer (ViT) based CLIP model trained by OFA-Sys for Chinese vision-language understanding. It supports zero-shot image classification, image-text retrieval, and related tasks using Chinese text. The model is available on Hugging Face with Transformers integration and is suitable for developers working with Chinese multimodal content.
Chinese Clip Vit Base Patch16 is an Other AI product. It focuses on performing zero-shot image classification and cross-modal retrieval using Chinese language prompts. It is built as an open-source project for developers. Chinese Clip Vit Base Patch16 is open source under the MIT license. It runs on the web and API.
Behind Chinese Clip Vit Base Patch16 is OFA-Sys, and the product first shipped in 2022. Development happens publicly on GitHub with 6k stars. Key capabilities include zero-shot classification, vision-language model, and chinese text support.
Latest indexed changes and source events
OFA-Sys/chinese-clip-vit-base-patch16 verified by the PulseGate indexer
Other apps tracked under the same category.