GLM-4.7-AWQ is a quantized variant of the GLM-4.7 open-source large language model, optimized using AWQ for reduced memory usage and faster inference on GPUs. It supports advanced features such as function/tool calling via a specialized chat template and can be used for text generation, coding assistance, and structured output tasks. The model is hosted on Hugging Face and is intended for developers who want to self-host or run capable LLMs locally or in their own infrastructure.
GLM sits in PulseGate's Quantised & converted weights category. It focuses on running large language models efficiently on consumer or edge hardware without sacrificing too much performance. GLM is an open-source project aimed at developers. The project is open source (Apache-2.0). It ships for the web, the command line, and API, and it can be self-hosted.
QuantTrio builds and maintains GLM, and it first shipped in 2025. The project is developed in the open on GitHub with 4.4k stars. Among its 4 catalogued features are quantized weights, tool calling, and chat template. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do