A 1B parameter version of Meta's Llama 3.2 Instruct model, quantized to 4-bit and packaged for the MLC-LLM runtime. It enables high-performance local inference across GPUs, CPUs, and mobile devices. The model is distributed via Hugging Face and supports standard chat templates for conversational use.
In the Other AI space, Llama 3.2 1B Instruct Q4f16 1 MLC takes a focused approach. It focuses on running efficient quantized large language models on local devices and various hardware backends. It is built as an open-source project for developers. Llama 3.2 1B Instruct Q4f16 1 MLC is open source under the Apache-2.0 license. It runs on the web, the command line, and API, and it can be self-hosted.
MLC AI builds and maintains Llama 3.2 1B Instruct Q4f16 1 MLC, and it first shipped in 2023. Development happens publicly on GitHub with 23k stars and 6 commits in the last 90 days. PulseGate's similarity index places it among 6 comparable projects. Among its 3 catalogued features are Quantized Weights, MLC Runtime, and Instruct Tuned.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do