Laguna-S-2.1-NVFP4 is a quantized version of the Laguna-S-2.1 Mixture-of-Experts language model optimized for efficient local inference. It uses NVFP4 (NVIDIA FP4) compression on the expert layers, enabling lower memory usage while maintaining strong text-generation performance. The model is available on Hugging Face for download and can be run via popular inference libraries such as transformers or vLLM.
Laguna S is a Foundation models & chat project. It focuses on running high-performance text generation models locally with 4-bit NVFP4 quantization without needing massive GPU memory. It is built as an open-source project for developers and AI researchers. Laguna S is open source under the MIT license. It ships for the web, the command line, and API.
It is developed by poolside, and it first shipped in 2023. Development happens publicly on GitHub with 121.4k stars and 1.2k commits in the last 90 days. Key capabilities include Text Generation, mixture of Experts, and 4-bit Quantization.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do