GLM-4.7-AWQ is a quantized variant of the GLM-4.7 open-source large language model, optimized using AWQ for reduced memory usage and faster inference on GPUs. It supports advanced features such as function/tool calling via a specialized chat template and can be used for text generation, coding assistance, and structured output tasks. The model is hosted on Hugging Face and is intended for developers who want to self-host or run capable LLMs locally or in their own infrastructure.
In the Foundation models & chat space, GLM takes a focused approach. It focuses on running large language models efficiently on consumer or edge hardware without sacrificing too much performance. GLM is an open-source project aimed at developers. The project is open source (Apache-2.0). GLM is available on the web, the command line, and API, and it can be self-hosted.
QuantTrio builds and maintains GLM, and the product first shipped in 2025. The project is developed in the open on GitHub with 4.4k stars. Among its 4 catalogued features are quantized weights, tool calling, and chat template. It exposes integrations via a public API.
Latest indexed changes and source events
QuantTrio/GLM-4.7-AWQ verified by the PulseGate indexer
Other apps tracked under the same category.