The Llama 3.2 1B Instruct Q4_K_M is a quantized GGUF version of a text generation model hosted on Hugging Face. It belongs to the class of foundation models and carries tags for conversational use, Llama architecture, and llama.cpp compatibility.
The repository is maintained under the hugging-quants organization and includes the file llama-3.2-1b-instruct-q4_k_m.gguf. It is marked as quantized, with a total file size of 807690656 bytes. The model card lists support for eight languages and references the original Meta Llama 3.2 1B Instruct base along with PyTorch and GGUF formats.
A chat template is defined in the repository data using special tokens such as bos_token, eos_token, and structured role headers for assistant and user messages. This enables direct use in compatible inference engines without additional conversion. The page indicates four discussions, two of which are closed.
No pricing, licensing terms, or specific target audience beyond the model tags are stated in the repository metadata.
In the Foundation models & chat space, Llama 3.2 1B Instruct Q4 K M takes a focused approach. It focuses on running lightweight instruction-tuned LLMs locally with minimal hardware requirements using GGUF format. Llama 3.2 1B Instruct Q4 K M is an open-source project aimed at developers running local AI. Llama 3.2 1B Instruct Q4 K M is open source under the MIT license. Llama 3.2 1B Instruct Q4 K M is available on the web, the command line, and API.
It is developed by Hugging Quants, and it first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 19 similar projects. Key capabilities include Quantized Model, Local Inference, and Instruction Following.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do