This Hugging Face model is a highly compressed 2-bit quantized variant of the Qwen3.6-35B-A3B large language model. It achieves approximately 100% quality retention compared to FP8 while drastically reducing memory requirements. The model is designed for local or edge inference and is suitable for developers building applications that require strong reasoning and coding capabilities from a smaller footprint. It includes custom tokenization and prompt formatting templates.
Qwen3.6 35B A3B Escha W2 is a Foundation models & chat project. It focuses on running large language models efficiently on limited hardware without significant quality loss. Qwen3.6 35B A3B Escha W2 is an open-source project aimed at AI researchers and developers. Qwen3.6 35B A3B Escha W2 is open source under the Open Source license. It ships for the web, the command line, and API.
EschaLabs builds and maintains Qwen3.6 35B A3B Escha W2, and it first shipped in 2025. Key capabilities include 2-bit Quantization, High Fidelity Retention, and Efficient Inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do