kimi-k2.6-eagle3-mla is an Eagle3 MTP draft model that incorporates Multi-Latent Attention for speculative decoding acceleration of the Kimi-K2.6 model. Hosted on Hugging Face by LightSeek Foundation, it serves as a specialized component in text generation pipelines that employ speculative decoding to improve inference speed.
The model was trained using TorchSpec, an online speculative decoding training framework that performs FSDP training and inference concurrently. Its use of Multi-Latent Attention provides two concrete advantages over an MHA draft model when paired with Kimi-K2.6: it consumes less KV cache memory, thereby lowering serving memory pressure, and it aligns with Kimi-K2.6's own MLA architecture for more natural integration into the inference engine's KV cache handling. The model card lists it under the speculative-decoding and eagle3 tags and identifies its primary task as text generation.
It can be deployed through vLLM or SGLang. The repository supplies instructions for launching a server with either backend and includes a citation section for academic reference. The model files are provided in Safetensors format.
The license is listed as other. No pricing information appears in the repository.
Kimi K2.6 Eagle3 Mla sits in PulseGate's Foundation models & chat category. It focuses on accelerating inference speed of the Kimi-K2.6 model using speculative decoding with reduced memory usage. It is built as an open-source project for AI inference engineers. Kimi K2.6 Eagle3 Mla is open source under the MIT license. It runs on the web and API.
Behind Kimi K2.6 Eagle3 Mla is lightseekorg, and the product first shipped in 2026. Development happens publicly on GitHub with 203 stars and 43 commits in the last 90 days.
Latest indexed changes and source events
lightseekorg/kimi-k2.6-eagle3-mla verified by the PulseGate indexer
Other apps tracked under the same category.