This is a quantized version of Meta's Llama 3.1 8B Instruct model using FP8 precision for the KV cache, optimized for AMD hardware. It is hosted on Hugging Face and can be used with the Transformers library for text generation and chat applications. The model is intended for developers seeking efficient local inference with reduced memory requirements.
In the Foundation models & chat space, Llama 3.1 8B Instruct FP8 KV takes a focused approach. It focuses on running efficient quantized large language models locally on AMD hardware. It is built as an open-source project for developers. Llama 3.1 8B Instruct FP8 KV is open source under the Open Source license. Llama 3.1 8B Instruct FP8 KV is available on the web, the command line, and API.
It is developed by AMD, and the product first shipped in 2024. Key capabilities include Image Classification, transformers, and pyTorch.
Latest indexed changes and source events
amd/Llama-3.1-8B-Instruct-FP8-KV verified by the PulseGate indexer
Other apps tracked under the same category.