KVzap-mlp-Qwen3-8B is a model from NVIDIA that implements a fast, adaptive KV cache pruning method called KVzap. It uses a lightweight MLP to predict importance scores for each KV pair and removes those below a threshold. This accelerates both prefilling and decoding phases of LLM inference while maintaining output quality. It is based on the Qwen3-8B architecture.
KVzap Mlp Qwen3 8B sits in PulseGate's Other AI category. It focuses on reducing memory and compute requirements of large language model inference by pruning unimportant KV cache entries. KVzap Mlp Qwen3 8B is an open-source project aimed at AI developers and researchers. The project is open source (Apache-2.0). The product ships for the web and API.
Behind KVzap Mlp Qwen3 8B is NVIDIA, and the product first shipped in 2024. The project is developed in the open on GitHub with 1.1k stars and 12 commits in the last 90 days. Among its 3 catalogued features are KV cache pruning, adaptive pruning, and transformers compatible.
Latest indexed changes and source events
nvidia/KVzap-mlp-Qwen3-8B verified by the PulseGate indexer
Other apps tracked under the same category.