This repository hosts an AWQ-quantized version of Meta's Llama 3 70B Instruct model, created by casperhansen. It enables efficient inference of this powerful instruction-tuned model on consumer or enterprise hardware with lower VRAM requirements. The model supports chat templates and is compatible with standard Hugging Face transformers and inference tools.
Llama 3 70b Instruct sits in PulseGate's Foundation models & chat category. It focuses on running large 70B parameter language models with reduced memory usage through quantization. It is built as an open-source project for AI developers and researchers. Llama 3 70b Instruct is open source under the Open Source license. The product ships for the web, the command line, and API.
casperhansen builds and maintains Llama 3 70b Instruct, and the product first shipped in 2024. It operates in a well-populated space: PulseGate tracks 8 similar tools.
Latest indexed changes and source events
casperhansen/llama-3-70b-instruct-awq verified by the PulseGate indexer
Other apps tracked under the same category.