Microsoft's Phi-3 Vision is a 128K-context multimodal model that can process both images and text. It supports visual question answering and detailed image description. The instruct-tuned version is available on Hugging Face and designed for efficient on-device or cloud deployment.
Phi 3 Vision 128k Instruct is a Multimodal & vision project. It focuses on understanding and reasoning over images and long text sequences in a compact model. It is built as an open-source project for AI developers and researchers. The project is open source (MIT). It ships for the web, the command line, and API.
Behind Phi 3 Vision 128k Instruct is Microsoft, based in the United States, and it first shipped in 2024. Development happens publicly on GitHub with 3.8k stars and 31 commits in the last 90 days.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do