MiniCPM-V-2 is an open-source vision-language model from OpenBMB designed for visual question answering and general multimodal tasks. It combines strong language understanding with visual processing capabilities and is distributed on Hugging Face with Transformers support. The model is optimized for efficiency while maintaining high performance on benchmarks like ScreenSpot-Pro, making it suitable for research and local deployment.
MiniCPM V 2 is a Foundation models & chat project. It focuses on enabling efficient on-device or local multimodal AI that can understand images and answer questions about them. It is built as an open-source project for AI researchers and developers. The project is open source (Apache-2.0). It ships for the web, the command line, and API.
Behind MiniCPM V 2 is OpenBMB, and it first shipped in 2024. The project is developed in the open on GitHub with 26k stars and 26 commits in the last 90 days. Key capabilities include Visual Question Answering, Multimodal Understanding, and Transformers Compatible.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match