A 1B parameter version of Meta's Llama 3.2 Instruct model, quantized to 4-bit and packaged for the MLC-LLM runtime. It enables high-performance local inference across GPUs, CPUs, and mobile devices. The model is distributed via Hugging Face and supports standard chat templates for conversational use.
In the Other AI space, Llama 3.2 1B Instruct Q4f16 1 MLC takes a focused approach. It focuses on running efficient quantized large language models on local devices and various hardware backends. It is built as an open-source project for developers. Llama 3.2 1B Instruct Q4f16 1 MLC is open source under the Apache-2.0 license. It runs on the web, the command line, and API, and it can be self-hosted.
Behind Llama 3.2 1B Instruct Q4f16 1 MLC is MLC AI, and the product first shipped in 2023. Development happens publicly on GitHub with 23k stars and 6 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 6 similar tools. Key capabilities include Quantized Weights, MLC Runtime, and Instruct Tuned.
Latest indexed changes and source events
mlc-ai/Llama-3.2-1B-Instruct-q4f16_1-MLC verified by the PulseGate indexer