Ovis1.6-Llama3.2-3B is an open multimodal model that integrates vision understanding with the Llama 3.2 3B language model. It supports image and text inputs, tool calling, and follows a chat template optimized for instruction following. The model is distributed on Hugging Face and is suitable for developers building lightweight multimodal applications that require both language reasoning and visual comprehension.
In the Foundation models & chat space, Ovis1.6 Llama3.2 3B takes a focused approach. It focuses on enabling smaller open-source models to understand both text and images for multimodal applications without relying on massive proprietary systems. It is built as an open-source project for Multimodal AI developers. Ovis1.6 Llama3.2 3B is open source under the Apache-2.0 license. It runs on the web, the command line, and API.
ATH-MaaS builds and maintains Ovis1.6 Llama3.2 3B, and the product first shipped in 2024. Development happens publicly on GitHub with 1.5k stars and 1 commits in the last 90 days. Key capabilities include multimodal input, vision-language understanding, and tool use support.
Latest indexed changes and source events
ATH-MaaS/Ovis1.6-Llama3.2-3B verified by the PulseGate indexer