MiniCPM-V-4 is an open multimodal language model for processing single images, multiple images, and video. It is distributed through Hugging Face and Transformers for developers and researchers deploying local vision-language inference, including on mobile devices.
In the Multimodal & vision space, MiniCPM V 4 takes a focused approach. It focuses on running multimodal image and video understanding locally on phones and other devices. It is built as an open-source project for AI developers and researchers. The project is open source (Apache-2.0). MiniCPM V 4 is available on the web, the command line, and API, and it can be self-hosted.
It is developed by OpenBMB, and it first shipped in 2024. The project is developed in the open on GitHub with 26.3k stars and 15 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 12 similar projects. Among its 6 catalogued features are image understanding, multi-image input, and video understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do