This repository contains an NVFP4 quantized variant of GLM-5.2 for optimized inference on NVIDIA hardware. It includes support for tool calling and advanced reasoning modes. The model follows a specific system prompt format and can be used with compatible inference engines.
GLM sits in PulseGate's Foundation models & chat category. It focuses on running the GLM-5.2 model efficiently on NVIDIA GPUs using specialized quantization. GLM is an open-source project aimed at developers and AI researchers. GLM is open source under the Apache-2.0 license. It runs on the web, the command line, and API.
Behind GLM is lukealonso, and it first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 355 commits in the last 90 days. Key capabilities include Large Language Model, NVFP4 Quantization, and Tool Use Support. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do