A quantized (w4a16) version of Meta's Llama 3.1 70B Instruct model provided by RedHatAI. It maintains the strong instruction-following and reasoning capabilities of the original while using 4-bit weights for more efficient inference. The model is compatible with standard Hugging Face and vLLM inference stacks.
In the Foundation models & chat space, Meta Llama 3.1 70B Instruct Quantized.w4a16 takes a focused approach. It focuses on running the large 70B Llama 3.1 model with significantly reduced memory requirements while maintaining performance. Meta Llama 3.1 70B Instruct Quantized.w4a16 is an open-source project aimed at developers. The project is open source (MIT). Meta Llama 3.1 70B Instruct Quantized.w4a16 is available on the web, the command line, and API.
Behind Meta Llama 3.1 70B Instruct Quantized.w4a16 is RedHatAI, based in the United States, and the product first shipped in 2023. The GitHub repository has been archived. PulseGate's similarity index places it among 9 comparable tools. Among its 4 catalogued features are text generation, instruction tuning, and quantized weights.
Latest indexed changes and source events
RedHatAI/Meta-Llama-3.1-70B-Instruct-quantized.w4a16 verified by the PulseGate indexer
⚠ Archived
Other apps tracked under the same category.