This repository contains an INT4 (w4a16) quantized version of Meta's Llama-3.1-8B-Instruct model. It enables efficient inference on hardware with limited resources while preserving most of the original model's instruction-following capabilities. The model is distributed for use with Transformers and compatible inference engines.
Meta Llama 3.1 8B Instruct Quantized.w4a16 is a Foundation models & chat product. It focuses on deploying the Llama 3.1 8B Instruct model with significantly reduced memory footprint for local or on-premise use. Meta Llama 3.1 8B Instruct Quantized.w4a16 is an open-source project aimed at developers and enterprises. The project is open source (MIT). The product ships for the web, the command line, and API.
It is developed by RedHatAI, and the product first shipped in 2023. The GitHub repository has been archived. PulseGate's similarity index places it among 10 comparable tools.
Latest indexed changes and source events
RedHatAI/Meta-Llama-3.1-8B-Instruct-quantized.w4a16 verified by the PulseGate indexer
⚠ Archived
Other apps tracked under the same category.