PulseGateLive intelligence on AI-era software
Coverage
180,820

Software tracked

 

Freshness
35 min ago

Last update

 

Cadence
516/day

7-day average

Indexed today: 469

PulseGate

Live intelligence on the software shipping in the AI era — apps, models, agents, and infra.

Software is shipping faster than ever, and a growing share of it lives outside the official app stores. PulseGate tracks it live — free, for builders, analysts, and everyone keeping up.

Follow
GitHubX (Twitter)LinkedIn

Platform

  • All Apps
  • Full Index
  • Categories
  • Industry Updates
  • Data Sources
  • Coverage Rules
  • Glossary
  • Embed Widget

Support

  • Help Center
  • Suggest a URL
  • Report an Issue

Company

  • About
  • Dimaxia
  • Press & Data
  • Contact
  • Platform Status

Legal

  • Privacy
  • Terms
  • Disclaimer

© 2026 PulseGate. Operated by Dimaxia, a brand of Dymaxio s.r.o., Prague, Czech Republic.·

All systems operational
PulseGate
KM
KVzap Mlp Qwen3 8B
Visit ↗
  1. Home/
  2. Other AI/
  3. KVzap Mlp Qwen3 8B
←Back to results
KM

KVzap Mlp Qwen3 8B

huggingface.co·Infrastructure

KVzap-mlp-Qwen3-8B is a model from NVIDIA that implements a fast, adaptive KV cache pruning method called KVzap. It uses a lightweight MLP to predict importance scores for each KV pair and removes those below a threshold. This accelerates both prefilling and decoding phases of LLM inference while maintaining output quality. It is based on the Qwen3-8B architecture.

Open SourceApache-2.0
WebAPI
K
KVzap Mlp Qwen3 8B preview
Visit huggingface.co↗
⭐1.1k
stars
🍴159
forks
✓3
features
📅2024
since

Overview

3 features

KVzap Mlp Qwen3 8B sits in PulseGate's Other AI category. It focuses on reducing memory and compute requirements of large language model inference by pruning unimportant KV cache entries. KVzap Mlp Qwen3 8B is an open-source project aimed at AI developers and researchers. The project is open source (Apache-2.0). The product ships for the web and API.

Behind KVzap Mlp Qwen3 8B is NVIDIA, and the product first shipped in 2024. The project is developed in the open on GitHub with 1.1k stars and 12 commits in the last 90 days. Among its 3 catalogued features are KV cache pruning, adaptive pruning, and transformers compatible.

  • ✓KV cache pruning
  • ✓Adaptive pruning
  • ✓Transformers compatible

Tags

kv-cache-pruningllm-inferenceqwen3

AI capabilities

Inference: LocalWeights: Open

Built with & integrations

Hosting
aws
Runs on
BrowserAPI-only

Trust & compliance

LicenseApache-2.0
Verified signals
✓ HTTPS✓ Privacy Policy✓ Terms of Service✓ Open Source✓ Free tier✓ GitHub · ★ 1.1k✓ Active maintenance
Legal
Privacy Policy →Terms of Service →

Recent events

Latest indexed changes and source events

  1. IndexedJul 21, 5:09 PM

    nvidia/KVzap-mlp-Qwen3-8B verified by the PulseGate indexer

    Source: PulseGate indexerOpen ↗

Frequently asked questions about KVzap Mlp Qwen3 8B

What is KVzap Mlp Qwen3 8B?
KVzap Mlp Qwen3 8B focuses on reducing memory and compute requirements of large language model inference by pruning unimportant KV cache entries. It is catalogued under Other AI on PulseGate.
Who should use KVzap Mlp Qwen3 8B?
KVzap Mlp Qwen3 8B is an open-source project built for AI developers and researchers.
Does KVzap Mlp Qwen3 8B have a free plan?
Yes — KVzap Mlp Qwen3 8B is open source under the Apache-2.0 license and free to use.
What platforms does KVzap Mlp Qwen3 8B run on?
KVzap Mlp Qwen3 8B runs on the web and API.
Is KVzap Mlp Qwen3 8B still active?
PulseGate's liveness checks currently classify KVzap Mlp Qwen3 8B as active. The GitHub repository shows 12 commits in the last 90 days.
What tools are similar to KVzap Mlp Qwen3 8B?
Similar tools tracked by PulseGate include 1 Fp8 Kv, Qwen3 8B, and Qwen3 30B A3B.1 Fp8 KvQwen3 8BQwen3 30B A3B
Who develops KVzap Mlp Qwen3 8B?
KVzap Mlp Qwen3 8B is developed by NVIDIA.
How long has KVzap Mlp Qwen3 8B been around?
KVzap Mlp Qwen3 8B first shipped in 2024.

At a glance

Pricing
Open Source
Platforms
Web
Languages
English
Open source
Yes · ★ 1.1k
License
Apache-2.0
First seen
Dec 18, 2025
Activity
🟢 Active
Status
🟢 Active
Built for
AI developers and researchers
Model
Open source
Solves
Reducing memory and compute requirements of large language model inference by pruning unimportant KV cache entries.

Developer

NVIDIA
Community-driven
↗ GitHub

Open source

View on GitHub →
⭐ Stars
1,143
🍴 Forks
159
Open issues
7
Last commit
3w ago
Commits 90d
12
Contributors
28
Authorship
Community-driven
Default branch
main
Latest release
v0.5.4 · 4w ago

Live coverage

Confidence
Medium · 70
Indexed
Jul 21, 2026
Lifecycle
Alive
Activity
Active
First seen
Dec 2025
Last seen
1w ago
Identity audit (10)
Entity ID
cmruwts250f0hqoarxrw3p2xd
Slug
nvidia-kvzap-mlp-qwen3-8b-huggingface-co
Lifecycle last checked
Jul 21, 2026
Verification state
Indexed for public listing
Claim / listing state
Unclaimed · listed: yes
Index status
Included in index
Latest evidence snapshot
Jul 21, 2026
Timeline basis
Indexed-at chronology (no inferred launch/funding milestones).
Last updated
Jul 24, 2026
Canonical URL
https://huggingface.co/nvidia/KVzap-mlp-Qwen3-8B

Similar apps

Other apps tracked under the same category.

  • 1 Fp8 Kv
    huggingface.co
  • Qwen3 8B
    huggingface.co
  • Qwen3 30B A3B
    huggingface.co
  • Qwen3 32B
    huggingface.co
  • Qwen3.5 122B A10B
    huggingface.co
  • Qwen3.6 27B
    huggingface.co