nvidia/KVzap-mlp-Llama-3.1-8B-Instruct is a variant of the Llama 3.1 8B model that incorporates KVzap, a method for pruning the key-value cache using a lightweight MLP to predict importance scores. This accelerates both prefilling and decoding phases of LLM inference while maintaining model quality. It is hosted on Hugging Face and can be used with the Transformers library.
KVzap Mlp Llama 3.1 8B Instruct is an Other AI project. It focuses on reducing the computational cost and memory usage of large language model inference through intelligent KV cache pruning. KVzap Mlp Llama 3.1 8B Instruct is an open-source project aimed at machine learning engineers and researchers. KVzap Mlp Llama 3.1 8B Instruct is open source under the Apache-2.0 license. KVzap Mlp Llama 3.1 8B Instruct is available on the web and API.
NVIDIA builds and maintains KVzap Mlp Llama 3.1 8B Instruct, and it first shipped in 2024. The project is developed in the open on GitHub with 1.2k stars and 14 commits in the last 90 days. Key capabilities include KV Cache Pruning, Fast Inference, and llama-based. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do