GLM-4V-9B is a multimodal variant of the GLM-4 large language model series. It combines a vision encoder with a 9-billion-parameter language model, enabling it to process images alongside text. The model supports both Chinese and English and is suitable for vision-language tasks such as image captioning, visual question answering, and document understanding.
In the Foundation models & chat space, Glm 4v 9b takes a focused approach. It focuses on providing an open multimodal model capable of understanding both images and text, with strong Chinese language performance. It is built as an open-source project for AI developers. Glm 4v 9b is open source under the Apache-2.0 license. It runs on the web and API.
ZAI-ORG builds and maintains Glm 4v 9b, and the product first shipped in 2024. Development happens publicly on GitHub with 7.1k stars. Key capabilities include Vision Language Model, multimodal, and Chinese English Support.
Latest indexed changes and source events
zai-org/glm-4v-9b verified by the PulseGate indexer
Other apps tracked under the same category.