This is a 4.75-bit PrismaQuant version of a Qwen3.6 35B-A3B model optimized for use with the vLLM inference engine. It supports multimodal inputs including vision and includes custom chat templates for tool use. The model is distributed on Hugging Face for efficient local or server-based deployment.
In the Foundation models & chat space, Qwen3.6 35B A3B PrismaQuant 4.75bit Vllm takes a focused approach. It focuses on deploying large Qwen language and vision models efficiently with very low bit-width quantization for reduced memory usage. Qwen3.6 35B A3B PrismaQuant 4.75bit Vllm is an open-source project aimed at developers. The project is open source (Open Source). The product ships for the web, API, and the command line.
rdtand builds and maintains Qwen3.6 35B A3B PrismaQuant 4.75bit Vllm, and the product first shipped in 2026. The project is developed in the open on GitHub with 97 stars and 603 commits in the last 90 days. PulseGate's similarity index places it among 14 comparable tools. Among its 3 catalogued features are Quantized Inference, Vision Support, and vLLM Compatibility.
Latest indexed changes and source events
rdtand/Qwen3.6-35B-A3B-PrismaQuant-4.75bit-vllm verified by the PulseGate indexer