Mage-VL is a multimodal foundation model hosted on Hugging Face. It accepts mixed visual inputs of images and video alongside text.
The model employs a chat template that counts image and video occurrences in a conversation. For each visual element it inserts dedicated vision tokens such as vision_start, image_pad or video_pad, and vision_end. A system prompt is added automatically when the first message is not from the system role. The template supports both string content and structured content lists that distinguish text from image or video objects.
It is delivered as a model repository under the microsoft organization on the Hugging Face platform. The repository supplies tokenizer configuration that defines special tokens including an image pad token, a video pad token, a vision start token, a vision end token, and an end-of-text token. Researchers and developers can download the model files and use the provided chat template for multimodal conversations.
No pricing, licensing terms, or additional capabilities are stated in the repository page.
Mage VL is a Foundation models & chat product. It focuses on understanding and reasoning over both images and video content using a single unified model. Mage VL is an open-source project aimed at AI researchers and developers. The project is open source (MIT). It runs on the web, the command line, and API.
Behind Mage VL is Microsoft, based in the United States, and the product first shipped in 2026. The project is developed in the open on GitHub with 673 stars and 32 commits in the last 90 days. Among its 5 catalogued features are multimodal understanding, image analysis, and video analysis.
Latest indexed changes and source events
Multimodal foundation model for image and video understanding from Microsoft verified by the PulseGate indexer
Other apps tracked under the same category.