This is a quantized 1B parameter Llama 3.2 model optimized with 8-bit weights using compressed-tensors. It enables efficient local inference for text generation tasks while maintaining strong performance. Hosted on Hugging Face, it is intended for developers integrating open-weight LLMs into applications with limited hardware resources.
Llama 3.2 1B Quantized.w8a8 is a Foundation models & chat project. It focuses on running large language models with reduced memory and compute requirements on local hardware. Llama 3.2 1B Quantized.w8a8 is an open-source project aimed at developers. The project is open source (Open Source). It ships for the web and API.
It is developed by RedHatAI, and it first shipped in 2024. PulseGate's similarity index places it among 7 comparable projects. Key capabilities include quantized weights, safetensors format, and 8-bit compression.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do