Xenova/clip-vit-base-patch16 is a model repository on Hugging Face that hosts a version of the CLIP ViT-B/16 architecture. It is provided for use in JavaScript-based machine learning environments through compatibility with Transformers.js and ONNX.js. The repository was created on 19 May 2023 and last modified on 8 October 2024. It has recorded over 868000 downloads to date.
The entry contains configuration details for a tokenizer that includes specific added tokens for beginning-of-sequence, end-of-sequence, padding, and unknown tokens, all set with defined stripping and normalization behaviors. No further capabilities, tasks, or technical specifications are described on the page itself. The model is listed without any assigned inference providers and with discussions enabled under recent sorting.
It forms part of the broader collection of models hosted on the Hugging Face platform, which focuses on open source and open science initiatives in artificial intelligence. The page provides no information on licensing, pricing, target audience, or deployment methods beyond the repository context.
In the Other AI space, Clip Vit Base Patch16 takes a focused approach. It focuses on running CLIP vision-language models directly in the browser or Node.js without Python dependencies. It is built as an open-source project for developers. Clip Vit Base Patch16 is open source under the Open Source license. Clip Vit Base Patch16 is available on the web, the command line, and API.
It is developed by Xenova. Among its 3 catalogued features are Image Embeddings, Text Embeddings, and Zero-shot Classification.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do