This is an AWQ (Activation-aware Weight Quantization) 4-bit quantized version of Meta's Llama 3 8B Instruct model, hosted on Hugging Face. It enables efficient local or server-based inference with significantly lower memory usage while preserving most of the original model's performance. The model includes a chat template and is compatible with popular inference engines like vLLM and Hugging Face Transformers. It is intended for developers building local AI applications or optimizing LLM deployment.
Llama 3 8b Instruct is a Foundation models & chat project. It focuses on running large language models efficiently on consumer or edge hardware with reduced memory and compute requirements. Llama 3 8b Instruct is an open-source project aimed at developers. Llama 3 8b Instruct is open source under the Open Source license. It ships for the web, the command line, and API.
Behind Llama 3 8b Instruct is casperhansen, and it first shipped in 2024. PulseGate's similarity index places it among 7 comparable projects. Key capabilities include quantized weights, chat template, and speculative decoding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do