Ovis1.6-Llama3.2-3B is an open multimodal model that integrates vision understanding with the Llama 3.2 3B language model. It supports image and text inputs, tool calling, and follows a chat template optimized for instruction following. The model is distributed on Hugging Face and is suitable for developers building lightweight multimodal applications that require both language reasoning and visual comprehension.
Ovis1.6 Llama3.2 3B sits in PulseGate's Multimodal & vision category. It focuses on enabling smaller open-source models to understand both text and images for multimodal applications without relying on massive proprietary systems. It is built as an open-source project for Multimodal AI developers. Ovis1.6 Llama3.2 3B is open source under the Apache-2.0 license. It ships for the web, the command line, and API.
ATH-MaaS builds and maintains Ovis1.6 Llama3.2 3B, and it first shipped in 2024. The project is developed in the open on GitHub with 1.5k stars and 1 commit in the last 90 days. Among its 3 catalogued features are multimodal input, vision-language understanding, and tool use support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do