Laguna-S-2.1-INT4 is a compressed, 4-bit quantized variant of Poolside's Laguna-S-2.1 Mixture-of-Experts language model. It uses compressed-tensors quantization to reduce model size and memory requirements while maintaining generation quality. The model is designed for text-generation tasks and can be loaded via the Hugging Face Transformers library with custom Laguna modeling code.
In the Foundation models & chat space, Laguna S takes a focused approach. It focuses on running efficient large language model inference on hardware with limited memory using INT4 quantization. It is built as an open-source project for developers and AI engineers. Laguna S is open source under the MIT license. Laguna S is available on the web, the command line, and API.
Poolside builds and maintains Laguna S, and it first shipped in 2023. Development happens publicly on GitHub with 121.6k stars and 1.2k commits in the last 90 days. Key capabilities include Text Generation, Quantized Inference, and mixture of Experts. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do