unsloth/GLM-4.7-Flash-FP8-Dynamic is an FP8 dynamically quantized version of the GLM-4.7-Flash model. Created by Unsloth, it enables efficient local inference with significantly reduced VRAM requirements while maintaining model quality. The model includes a custom chat template and is compatible with the Transformers library and other inference engines.
GLM 4.7 Flash FP8 Dynamic is a Foundation models & chat product. It focuses on running large language models with reduced memory and faster inference using FP8 quantization. GLM 4.7 Flash FP8 Dynamic is an open-source project aimed at developers and researchers. The project is open source (Apache-2.0). The product ships for the web, the command line, and API.
Unsloth builds and maintains GLM 4.7 Flash FP8 Dynamic, and the product first shipped in 2023. The project is developed in the open on GitHub with 68.4k stars and 1.2k commits in the last 90 days.
Latest indexed changes and source events
unsloth/GLM-4.7-Flash-FP8-Dynamic verified by the PulseGate indexer
Other apps tracked under the same category.