This is a 6-bit quantized version of the GLM-4.7-Flash model, prepared by the LM Studio community for efficient inference using the MLX framework on Apple devices. It includes support for tool calling and follows a standard chat template. The model is hosted on Hugging Face for easy local deployment.
GLM 4.7 Flash sits in PulseGate's Foundation models & chat category. It focuses on running a capable GLM-4 model efficiently on local Apple hardware through quantization and MLX optimization. It is built as an open-source project for developers. GLM 4.7 Flash is open source under the MIT license. GLM 4.7 Flash is available on the web, the command line, and API.
lmstudio-community builds and maintains GLM 4.7 Flash, and the product first shipped in 2023. Development happens publicly on GitHub with 6.4k stars and 21 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 5 similar tools. Key capabilities include Large Language Model, Tool Calling, and Quantized Inference.
Latest indexed changes and source events
lmstudio-community/GLM-4.7-Flash-MLX-6bit verified by the PulseGate indexer
Other apps tracked under the same category.