This repository contains an INT4 (w4a16) quantized version of Meta's Llama-3.1-8B-Instruct model. It enables efficient inference on hardware with limited resources while preserving most of the original model's instruction-following capabilities. The model is distributed for use with Transformers and compatible inference engines.
Meta Llama 3.1 8B Instruct Quantized.w4a16 is a Text generation project. It focuses on deploying the Llama 3.1 8B Instruct model with significantly reduced memory footprint for local or on-premise use. Meta Llama 3.1 8B Instruct Quantized.w4a16 is an open-source project aimed at developers and enterprises. The project is open source (MIT). It runs on the web, the command line, and API.
Behind Meta Llama 3.1 8B Instruct Quantized.w4a16 is RedHatAI, and it first shipped in 2023. The GitHub repository has been archived. PulseGate's similarity index places it among 10 comparable projects.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do