Microsoft's Phi-3 Vision is a 128K-context multimodal model that can process both images and text. It supports visual question answering and detailed image description. The instruct-tuned version is available on Hugging Face and designed for efficient on-device or cloud deployment.
Phi 3 Vision 128k Instruct is a Foundation models & chat product. It focuses on understanding and reasoning over images and long text sequences in a compact model. It is built as an open-source project for AI developers and researchers. Phi 3 Vision 128k Instruct is open source under the MIT license. Phi 3 Vision 128k Instruct is available on the web, the command line, and API.
Behind Phi 3 Vision 128k Instruct is Microsoft, based in the United States, and the product first shipped in 2024. Development happens publicly on GitHub with 3.8k stars and 31 commits in the last 90 days.
Latest indexed changes and source events
microsoft/Phi-3-vision-128k-instruct verified by the PulseGate indexer