This repository contains an NVFP4 quantized variant of GLM-5.2 for optimized inference on NVIDIA hardware. It includes support for tool calling and advanced reasoning modes. The model follows a specific system prompt format and can be used with compatible inference engines.
GLM sits in PulseGate's Foundation models & chat category. It focuses on running the GLM-5.2 model efficiently on NVIDIA GPUs using specialized quantization. GLM is an open-source project aimed at developers and AI researchers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
It is developed by lukealonso, and the product first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 355 commits in the last 90 days. Among its 3 catalogued features are Large Language Model, NVFP4 Quantization, and Tool Use Support. It exposes integrations via a public API.
Latest indexed changes and source events
lukealonso/GLM-5.2-NVFP4 verified by the PulseGate indexer
Other apps tracked under the same category.