A quantized (w4a16) version of Meta's Llama 3.1 70B Instruct model provided by RedHatAI. It maintains the strong instruction-following and reasoning capabilities of the original while using 4-bit weights for more efficient inference. The model is compatible with standard Hugging Face and vLLM inference stacks.
Meta Llama 3.1 70B Instruct Quantized.w4a16 sits in PulseGate's Text generation category. It focuses on running the large 70B Llama 3.1 model with significantly reduced memory requirements while maintaining performance. Meta Llama 3.1 70B Instruct Quantized.w4a16 is an open-source project aimed at developers. Meta Llama 3.1 70B Instruct Quantized.w4a16 is open source under the MIT license. Meta Llama 3.1 70B Instruct Quantized.w4a16 is available on the web, the command line, and API.
Behind Meta Llama 3.1 70B Instruct Quantized.w4a16 is RedHatAI, based in the United States, and it first shipped in 2023. The GitHub repository has been archived. PulseGate's similarity index places it among 9 comparable projects. Key capabilities include text generation, instruction tuning, and quantized weights.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do