This repository provides GGUF quantized files for the Qwen3.6-35B-A3B model, enabling efficient CPU and GPU inference using tools like llama.cpp. The model supports multimodal inputs including vision and offers strong reasoning capabilities. It is designed for users who want to run powerful language models locally without relying on cloud APIs.
In the Foundation models & chat space, Qwen3.6 35B A3B takes a focused approach. It focuses on running large 35B parameter models efficiently on consumer hardware using quantized GGUF format. It is built as an open-source project for developers. Qwen3.6 35B A3B is open source under the Open Source license. The product ships for the web and the command line.
It is developed by ggml-org, and the product first shipped in 2026. Development happens publicly on GitHub with 97 commits in the last 90 days. The category is crowded — PulseGate's index counts 25 comparable apps. Key capabilities include GGUF Format, quantized, and Local Inference.
Latest indexed changes and source events
ggml-org/Qwen3.6-35B-A3B-GGUF verified by the PulseGate indexer
Other apps tracked under the same category.