This is an AWQ (INT4) quantized version of Meta's Llama 3.1 70B Instruct model, optimized for reduced memory usage while maintaining performance. It includes a detailed chat template and is suitable for local inference. The model is provided by the hugging-quants organization on Hugging Face.
Meta Llama 3.1 70B Instruct sits in PulseGate's Text generation category. It focuses on running large 70B language models efficiently on consumer or enterprise hardware with reduced memory requirements. It is built as an open-source project for AI developers and researchers. The project is open source (MIT). It runs on the command line and API.
It is developed by hugging-quants, and it first shipped in 2023. The GitHub repository has been archived. It operates in a well-populated space: PulseGate tracks 13 similar projects. Among its 3 catalogued features are Quantized Weights, Instruction Tuned, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do