This repository hosts an AWQ-quantized version of Meta's Llama 3 70B Instruct model, created by casperhansen. It enables efficient inference of this powerful instruction-tuned model on consumer or enterprise hardware with lower VRAM requirements. The model supports chat templates and is compatible with standard Hugging Face transformers and inference tools.
Llama 3 70b Instruct is a Text generation project. It focuses on running large 70B parameter language models with reduced memory usage through quantization. Llama 3 70b Instruct is an open-source project aimed at AI developers and researchers. The project is open source (Open Source). It runs on the web, the command line, and API.
casperhansen builds and maintains Llama 3 70b Instruct. PulseGate's similarity index places it among 8 comparable projects.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do