This is a W4A16 quantized version of the Gemma-3-12B instruction-tuned (it) model created by abhishekchohan. It enables efficient local inference of the 12-billion parameter model on hardware with limited VRAM. The model includes a custom chat template optimized for conversational use and is compatible with the Hugging Face Transformers library.
In the Foundation models & chat space, Gemma 3 12b It Quantized W4A16 takes a focused approach. It focuses on running the Gemma 3 12B model on consumer hardware with reduced memory usage through 4-bit quantization. It is built as an open-source project for developers and researchers. The project is open source (Open Source). Gemma 3 12b It Quantized W4A16 is available on the web, the command line, and API.
abhishekchohan builds and maintains Gemma 3 12b It Quantized W4A16. It competes in a saturated segment with 25 similar projects in PulseGate's index. Among its 3 catalogued features are Instruction Tuning, Quantized Inference, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do