This is a quantized version of Meta's Llama 3.1 8B Instruct model using FP8 precision for the KV cache, optimized for AMD hardware. It is hosted on Hugging Face and can be used with the Transformers library for text generation and chat applications. The model is intended for developers seeking efficient local inference with reduced memory requirements.
Llama 3.1 8B Instruct FP8 KV sits in PulseGate's Text generation category. It focuses on running efficient quantized large language models locally on AMD hardware. It is built as an open-source project for developers. The project is open source (Open Source). It ships for the web, the command line, and API.
Behind Llama 3.1 8B Instruct FP8 KV is AMD, and it first shipped in 2024. Among its 5 catalogued features are Image Classification, transformers, and pyTorch.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do