GLM-4.7-Flash-MLX-8bit is a community-quantized version of the GLM-4 large language model using 8-bit precision and optimized for the MLX framework on Apple silicon. It supports tool calling and can be used locally through libraries such as Transformers or MLX. The model is hosted on Hugging Face for easy download and integration into local AI applications.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on running large language models efficiently on local Apple hardware. It is built as an open-source project for developers. GLM 4.7 Flash is open source under the MIT license. The product ships for the web, the command line, and API.
It is developed by lmstudio-community, and the product first shipped in 2023. Development happens publicly on GitHub with 6.4k stars and 21 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 6 similar tools. Key capabilities include 8-bit Quantization, MLX Optimization, and Tool Calling.
Latest indexed changes and source events
lmstudio-community/GLM-4.7-Flash-MLX-8bit verified by the PulseGate indexer
Other apps tracked under the same category.