GLM-4.7-Flash-MLX-8bit is a community-quantized version of the GLM-4 large language model using 8-bit precision and optimized for the MLX framework on Apple silicon. It supports tool calling and can be used locally through libraries such as Transformers or MLX. The model is hosted on Hugging Face for easy download and integration into local AI applications.
In the Foundation models & chat space, GLM 4.7 Flash takes a focused approach. It focuses on running large language models efficiently on local Apple hardware. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (MIT). GLM 4.7 Flash is available on the web, the command line, and API.
lmstudio-community builds and maintains GLM 4.7 Flash, and it first shipped in 2023. The project is developed in the open on GitHub with 6.4k stars and 21 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 6 similar projects. Key capabilities include 8-bit Quantization, MLX Optimization, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do