EAGLE3 is a speculative decoding method that extrapolates hidden states from LLMs to achieve significant speedups in token generation. This 8B model is based on Llama 3.1 Instruct and implements the EAGLE-3 technique. It has been shown to provide 2-3x faster inference compared to standard decoding while preserving output quality. The approach is trainable and compatible with various inference engines.
EAGLE3 LLaMA3.1 Instruct 8B sits in PulseGate's Foundation models & chat category. It focuses on accelerating the generation speed of large language models while maintaining output distribution consistency. It is built as an open-source project for developers. EAGLE3 LLaMA3.1 Instruct 8B is open source under the Open Source license. It runs on the web and API.
yuhuili builds and maintains EAGLE3 LLaMA3.1 Instruct 8B, and the product first shipped in 2023. Development happens publicly on GitHub with 2.5k stars.
Latest indexed changes and source events
yuhuili/EAGLE3-LLaMA3.1-Instruct-8B verified by the PulseGate indexer
Other apps tracked under the same category.