flash-attn3 is a collection of optimized GPU kernels implementing Flash Attention 3. It provides functions such as flash_attn_func, flash_attn_qkvpacked_func, and flash_attn_combine. The library is designed to be used with the kernels package and is targeted at developers optimizing large language model performance.
Flash Attn3 is an Other AI project. It focuses on accelerating attention mechanisms in transformer models through highly optimized GPU kernels for faster training and inference. Flash Attn3 is an open-source project aimed at developers. The project is open source (Apache-2.0). It ships for the web, the command line, and API.
It is developed by kernels-community, and it first shipped in 2024. The project is developed in the open on GitHub with 715 stars and 166 commits in the last 90 days. Among its 3 catalogued features are Flash Attention, GPU Kernels, and Optimized Inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match