GLM-4.5-Air-FP8 is an FP8 quantized variant of the GLM-4.5 model, designed for efficient local inference. It supports tool calling and includes a comprehensive chat template. The model is suitable for developers building local AI applications and is distributed via Hugging Face for use with standard inference frameworks.
GLM 4.5 Air is a Coding AI & assistants project. It focuses on running large GLM models locally with reduced memory footprint using FP8 quantization. It is built as an open-source project for developers. GLM 4.5 Air is open source under the Apache-2.0 license. It ships for the web and the command line, and it can be self-hosted.
Behind GLM 4.5 Air is zai-org, and it first shipped in 2025. The project is developed in the open on GitHub with 4.4k stars. Key capabilities include quantized model, tool calling, and chat template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do