Llama 3.2 1B Instruct Q8 0 is a quantized GGUF version of the Llama 3.2 1B Instruct model hosted on Hugging Face by the Hugging Quants organization. It belongs to the class of foundation models and supports text generation and conversational tasks across eight languages. The model card indicates it uses the Llama architecture from Meta and includes a specific chat template with tokens such as bos_token, eos_token, and structured role-based formatting for assistant responses.
The repository provides the file llama-3.2-1b-instruct-q8_0.gguf along with associated metadata confirming it is quantized. Total file size is listed as 1,321,079,200 bytes. It carries tags for GGUF, PyTorch, llama.cpp, and facebook/meta/llama families, positioning it for use in compatible inference runtimes that handle this format.
Hugging Quants maintains the repository, which has accumulated around 50 likes and hosts four discussions. The page forms part of the broader Hugging Face ecosystem for models, datasets, and related resources focused on open-source AI.
Llama 3.2 1B Instruct Q8 0 sits in PulseGate's Text generation category. It focuses on running the Llama 3.2 1B model efficiently on CPUs and lower-end hardware. It is built as an open-source project for developers. Llama 3.2 1B Instruct Q8 0 is open source under the MIT license. It runs on the web, the command line, and API.
Hugging Quants builds and maintains Llama 3.2 1B Instruct Q8 0, and it first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. The category is crowded — PulseGate's index counts 25 comparable projects. Key capabilities include GGUF Quantization, Local Inference, and Instruct Model.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do