Llama 3.2 1B Instruct Q8 0 is a quantized GGUF version of the Llama 3.2 1B Instruct model hosted on Hugging Face by the Hugging Quants organization. It belongs to the class of foundation models and supports text generation and conversational tasks across eight languages. The model card indicates it uses the Llama architecture from Meta and includes a specific chat template with tokens such as bos_token, eos_token, and structured role-based formatting for assistant responses.
The repository provides the file llama-3.2-1b-instruct-q8_0.gguf along with associated metadata confirming it is quantized. Total file size is listed as 1,321,079,200 bytes. It carries tags for GGUF, PyTorch, llama.cpp, and facebook/meta/llama families, positioning it for use in compatible inference runtimes that handle this format.
Hugging Quants maintains the repository, which has accumulated around 50 likes and hosts four discussions. The page forms part of the broader Hugging Face ecosystem for models, datasets, and related resources focused on open-source AI.
Llama 3.2 1B Instruct Q8 0 is a Foundation models & chat product. It focuses on running the Llama 3.2 1B model efficiently on CPUs and lower-end hardware. Llama 3.2 1B Instruct Q8 0 is an open-source project aimed at developers. The project is open source (MIT). It runs on the web, the command line, and API.
Behind Llama 3.2 1B Instruct Q8 0 is Hugging Quants, and the product first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. It competes in a saturated segment with 25 similar apps in PulseGate's index. Among its 3 catalogued features are GGUF Quantization, Local Inference, and Instruct Model.
Latest indexed changes and source events
hugging-quants/Llama-3.2-1B-Instruct-Q8_0-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.