This is a 4-bit quantized version of a Qwen1.5 Mixture-of-Experts model with 2.7B active parameters, optimized for chat use cases. It is hosted on Hugging Face and can be used with the Transformers library or local inference engines. The model is intended for developers integrating efficient language generation into applications.
In the Foundation models & chat space, Qwen1.5 MoE A2.7B Chat Quantized.w4a16 takes a focused approach. It focuses on running large language models with reduced memory and compute requirements on local hardware. Qwen1.5 MoE A2.7B Chat Quantized.w4a16 is an open-source project aimed at developers. The project is open source (Open Source). The product ships for the web and API.
It is developed by nm-testing, and the product first shipped in 2025. Among its 3 catalogued features are quantized weights, mixture of Experts, and chat template.
Latest indexed changes and source events
nm-testing/Qwen1.5-MoE-A2.7B-Chat-quantized.w4a16 verified by the PulseGate indexer
Other apps tracked under the same category.