The Llama 3.2 1B Instruct Q4_K_M is a quantized GGUF version of a text generation model hosted on Hugging Face. It belongs to the class of foundation models and carries tags for conversational use, Llama architecture, and llama.cpp compatibility.
The repository is maintained under the hugging-quants organization and includes the file llama-3.2-1b-instruct-q4_k_m.gguf. It is marked as quantized, with a total file size of 807690656 bytes. The model card lists support for eight languages and references the original Meta Llama 3.2 1B Instruct base along with PyTorch and GGUF formats.
A chat template is defined in the repository data using special tokens such as bos_token, eos_token, and structured role headers for assistant and user messages. This enables direct use in compatible inference engines without additional conversion. The page indicates four discussions, two of which are closed.
No pricing, licensing terms, or specific target audience beyond the model tags are stated in the repository metadata.
Llama 3.2 1B Instruct Q4 K M is a Foundation models & chat product. It focuses on running lightweight instruction-tuned LLMs locally with minimal hardware requirements using GGUF format. Llama 3.2 1B Instruct Q4 K M is an open-source project aimed at developers running local AI. The project is open source (MIT). It runs on the web, the command line, and API.
It is developed by Hugging Quants, and the product first shipped in 2023. The project is developed in the open on GitHub with 121.2k stars and 1.2k commits in the last 90 days. PulseGate's similarity index places it among 19 comparable tools. Among its 3 catalogued features are Quantized Model, Local Inference, and Instruction Following.
Latest indexed changes and source events
hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.