This is a GGUF-quantized variant of the Llama-3-8B-Instruct model with 32k context length, published on Hugging Face. It provides multiple quantization levels (Q2_K through IQ4_XS) that allow users to run the model locally with reduced memory requirements while maintaining instruction-following capabilities. Primarily used by developers and researchers who want to deploy or experiment with long-context LLMs on personal machines or self-hosted environments.
Llama 3 8B Instruct 32k sits in PulseGate's Foundation models & chat category. It focuses on running a long-context Llama 3 8B model efficiently on consumer hardware without cloud dependency. Llama 3 8B Instruct 32k is an open-source project aimed at developers. Llama 3 8B Instruct 32k is open source under the MIT license. Llama 3 8B Instruct 32k is available on the web and the command line, and it can be self-hosted.
Behind Llama 3 8B Instruct 32k is MaziyarPanahi, and it first shipped in 2023. Development happens publicly on GitHub with 122.1k stars and 1.2k commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 12 similar projects. Among its 4 catalogued features are GGUF Quantization, Extended Context, and Instruct Tuned.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do