GLM-4.5-Air-FP8 is an FP8 quantized variant of the GLM-4.5 model, designed for efficient local inference. It supports tool calling and includes a comprehensive chat template. The model is suitable for developers building local AI applications and is distributed via Hugging Face for use with standard inference frameworks.
In the Coding AI & assistants space, GLM 4.5 Air takes a focused approach. It focuses on running large GLM models locally with reduced memory footprint using FP8 quantization. It is built as an open-source project for developers. GLM 4.5 Air is open source under the Apache-2.0 license. It runs on the web and the command line, and it can be self-hosted.
Behind GLM 4.5 Air is zai-org, and the product first shipped in 2025. Development happens publicly on GitHub with 4.4k stars. Key capabilities include quantized model, tool calling, and chat template.
Latest indexed changes and source events
zai-org/GLM-4.5-Air-FP8 verified by the PulseGate indexer
Other apps tracked under the same category.