llama-3.3-70b-instruct-awq is a quantized version of Meta's Llama 3.3 70B Instruct model using the AWQ quantization method. This allows the large model to run with reduced memory requirements while maintaining performance. It includes a chat template optimized for instruction following and is compatible with Transformers and other inference frameworks.
Llama 3.3 70b Instruct sits in PulseGate's Foundation models & chat category. It focuses on running the large 70B Llama 3.3 model efficiently on consumer or enterprise hardware through quantization. It is built as an open-source project for developers and researchers running large language models locally. The project is open source (MIT). It ships for the web, the command line, and API.
Behind Llama 3.3 70b Instruct is casperhansen, and it first shipped in 2023. The GitHub repository has been archived. PulseGate's similarity index places it among 6 comparable projects. Among its 3 catalogued features are quantized weights, 70B parameters, and instruction tuned.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do