This is an 8-bit quantized version of the GLM-4.6V-Flash model using MLX, designed for efficient local execution. It supports multimodal inputs including text and images and includes tooling capabilities. The model is distributed via Hugging Face for use with local inference tools like LM Studio.
GLM 4.6V Flash is a Foundation models & chat product. It focuses on running powerful multimodal AI models efficiently on consumer hardware with reduced memory usage. GLM 4.6V Flash is an open-source project aimed at developers. The project is open source (MIT). GLM 4.6V Flash is available on the web, the command line, and API.
Behind GLM 4.6V Flash is lmstudio-community, and the product first shipped in 2024. The project is developed in the open on GitHub with 5.2k stars and 419 commits in the last 90 days. Among its 3 catalogued features are Vision Language Model, Quantized Inference, and Tool Calling.
Latest indexed changes and source events
lmstudio-community/GLM-4.6V-Flash-MLX-8bit verified by the PulseGate indexer
Other apps tracked under the same category.