An NVIDIA-optimized version of the GLM-5 large language model using NVFP4 precision for accelerated inference on compatible GPUs. It includes support for tool calling and follows standard chat templates, making it suitable for deployment in performance-sensitive local or on-premise environments.
In the Foundation models & chat space, GLM 5 takes a focused approach. It focuses on running large GLM models efficiently on NVIDIA hardware with optimized precision formats. GLM 5 is an open-source project aimed at developers. The project is open source (Apache-2.0). GLM 5 is available on the web and the command line, and it can be self-hosted.
It is developed by nvidia, and the product first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 354 commits in the last 90 days.
Latest indexed changes and source events
nvidia/GLM-5-NVFP4 verified by the PulseGate indexer
Other apps tracked under the same category.