This is a W4A16 quantized version of the Gemma-3-12B instruction-tuned (it) model created by abhishekchohan. It enables efficient local inference of the 12-billion parameter model on hardware with limited VRAM. The model includes a custom chat template optimized for conversational use and is compatible with the Hugging Face Transformers library.
In the Foundation models & chat space, Gemma 3 12b It Quantized W4A16 takes a focused approach. It focuses on running the Gemma 3 12B model on consumer hardware with reduced memory usage through 4-bit quantization. Gemma 3 12b It Quantized W4A16 is an open-source project aimed at developers and researchers. The project is open source (Open Source). It runs on the web, the command line, and API.
It is developed by abhishekchohan, and the product first shipped in 2025. It competes in a saturated segment with 25 similar apps in PulseGate's index. Among its 3 catalogued features are Instruction Tuning, Quantized Inference, and Chat Template.
Latest indexed changes and source events
abhishekchohan/gemma-3-12b-it-quantized-W4A16 verified by the PulseGate indexer
Other apps tracked under the same category.