This is an 8-bit quantized version of the GLM-4.6V-Flash model using MLX, designed for efficient local execution. It supports multimodal inputs including text and images and includes tooling capabilities. The model is distributed via Hugging Face for use with local inference tools like LM Studio.
GLM 4.6V Flash is a Multimodal & vision project. It focuses on running powerful multimodal AI models efficiently on consumer hardware with reduced memory usage. It is built as an open-source project for developers. GLM 4.6V Flash is open source under the MIT license. It ships for the web, the command line, and API.
It is developed by lmstudio-community, and it first shipped in 2024. The project is developed in the open on GitHub with 5.2k stars and 419 commits in the last 90 days. Among its 3 catalogued features are Vision Language Model, Quantized Inference, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do