This is a GPTQ-Int4 quantized version of Alibaba's Qwen 3.5 35B (with A3B architecture) model hosted on Hugging Face. It supports text and vision inputs and is designed for efficient local or self-hosted inference using libraries such as transformers, vLLM, or llama.cpp. The model is suitable for developers who need strong language understanding and generation capabilities without requiring high-end hardware. It includes support for tool use and advanced prompting techniques.
In the Foundation models & chat space, Qwen3.5 35B A3B takes a focused approach. It focuses on running a large multimodal language model locally with significantly reduced VRAM usage through quantization. It is built as an open-source project for developers. Qwen3.5 35B A3B is open source under the Apache-2.0 license. It ships for the web, the command line, and API.
Behind Qwen3.5 35B A3B is Qwen, and it first shipped in 2023. The project is developed in the open on GitHub with 31k stars and 3.7k commits in the last 90 days. The category is crowded — PulseGate's index counts 25 comparable apps. Among its 3 catalogued features are Quantized LLM, multimodal support, and text generation.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do