This is a GGUF-quantized distribution of Meta's Llama 3.2 1B Instruct model hosted on Hugging Face. It enables efficient local inference on CPUs and GPUs using tools such as llama.cpp, Ollama, and LM Studio. The repository provides multiple quantization levels for different hardware and performance trade-offs, making small but capable language models accessible for on-device and edge use cases.
Llama 3.2 1B Instruct is a Foundation models & chat project. It focuses on running a capable 1B-parameter language model locally with low resource usage. Llama 3.2 1B Instruct is an open-source project aimed at developers. The project is open source (MIT). It ships for the web, the command line, and API, and it can be self-hosted.
It is developed by bartowski, and it first shipped in 2023. Development happens publicly on GitHub with 121.3k stars and 1.2k commits in the last 90 days. It competes in a saturated segment with 25 similar apps in PulseGate's index. Key capabilities include Multiple Quantizations, CPU Inference, and Local Execution.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do