This repository hosts an AWQ 4-bit quantized version of the GLM-4.7-Flash model, optimized for reduced memory usage and faster inference. It retains tool-calling capabilities and a full chat template, making it suitable for local or hosted deployment where full-precision models would be too large. The model is provided for developers seeking efficient open-weight alternatives.
GLM 4.7 Flash is a Foundation models & chat product. It focuses on running the GLM-4.7-Flash model efficiently on consumer hardware through 4-bit quantization. GLM 4.7 Flash is an open-source project aimed at developers. The project is open source (Apache-2.0). GLM 4.7 Flash is available on the web, the command line, and API.
Behind GLM 4.7 Flash is cyankiwi, and the product first shipped in 2025. The project is developed in the open on GitHub with 4.4k stars. Among its 3 catalogued features are quantized weights, tool calling, and chat template.
Latest indexed changes and source events
cyankiwi/GLM-4.7-Flash-AWQ-4bit verified by the PulseGate indexer
Other apps tracked under the same category.