Llama 3.2 11B Vision Instruct is an open vision-language model from Meta that can process both text and images. It supports instruction following across multimodal inputs and is distributed on Hugging Face for local and cloud inference. The model is suitable for developers building applications that require visual understanding combined with natural language generation.
In the Foundation models & chat space, Llama 3.2 11B Vision Instruct takes a focused approach. It focuses on enabling vision-language tasks such as image captioning, visual question answering, and multimodal reasoning in a single open model. It is built as an open-source project for developers. Llama 3.2 11B Vision Instruct is open source under the Open Source license. It runs on the web, the command line, and API.
Behind Llama 3.2 11B Vision Instruct is Meta, based in the United States, and the product first shipped in 2024. Development happens publicly on GitHub with 7.7k stars. It operates in a well-populated space: PulseGate tracks 5 similar tools. Key capabilities include Vision Understanding, Instruction Following, and Multimodal Input.
Latest indexed changes and source events
meta-llama/Llama-3.2-11B-Vision-Instruct verified by the PulseGate indexer