Skip to content
Alternatives
Software like Qwen3.5 35B A3B
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- Qwen3 30B A3Bhuggingface.coQwen3-30B-A3B-GPTQ-Int4 is a GPTQ 4-bit quantized variant of the Qwen3 30B-A3B model. It supports advanced capabilities such as tool calling and follows a detailed chat template for conversational use. The model is optimized for local or server-based inference using tools like llama.cpp or Hugging Face Text Generation Inference while maintaining strong performance.
- Qwen3.5 27Bhuggingface.coThis is an INT4 GPTQ-quantized version of the Qwen3.5-27B model from the Qwen team. It supports multimodal inputs including text, images, and video. The quantization enables more efficient inference while maintaining strong performance across various tasks.
- Qwen3.5 122B A10Bhuggingface.coThis is a GPTQ-Int4 quantized variant of Alibaba's Qwen3.5-122B-A10B model. It supports multimodal inputs (including vision and video) and advanced features such as tool calling. The model is distributed on Hugging Face and is intended for developers who want to run a high-performance open model locally or on modest GPU hardware.
- Qwen3.6 35B A3Bhuggingface.coThis is a GPTQ Int4 quantized version of the Qwen3.6-35B-A3B model optimized for reduced memory usage while maintaining performance. It supports vision inputs, complex chat templates, and multimodal content. The model is hosted on Hugging Face for use with Transformers and inference engines.
- Qwen3.5 35B A3Bhuggingface.coA 4-bit AWQ quantized checkpoint of the Qwen3.5-35B-A3B Mixture-of-Experts model. It supports multimodal inputs including images and video and is optimized for local execution using tools compatible with the GGUF or AWQ format. The model is intended for developers wanting high-performance local AI capabilities.
- Qwen3.6 35B A3Bhuggingface.coQwen3.6-35B-A3B-AWQ-4bit is a quantized version of a large language model, designed for efficient inference on local hardware. It is suitable for developers and researchers who need high-performance language models with reduced resource requirements.
- Qwen3.5 27Bhuggingface.coThis is an AWQ-quantized 27B parameter version of Alibaba's Qwen3.5 language model. It supports text generation, tool use, and multimodal inputs while requiring significantly less memory than the original. The model is distributed via Hugging Face for use with popular inference frameworks.
- Qwen3 14Bhuggingface.coA GPTQ 4-bit quantized version of the Qwen3-14B model, optimized for lower memory usage while maintaining strong performance. It includes support for tool calling and follows the Qwen chat template. Suitable for local inference on GPUs with limited VRAM using libraries such as transformers or vLLM.
- Qwen3.5 4Bhuggingface.coThis is an AWQ-quantized 4B parameter version of the Qwen 3.5 model optimized for efficient inference. It supports advanced features including tool use and vision capabilities. The model is distributed on Hugging Face for use with Transformers or local inference engines.
- Qwen3 30B A3Bhuggingface.coQwen3-30B-A3B-GGUF provides GGUF quantized weights for the Qwen3 30B-A3B model, enabling efficient local execution with tools like llama.cpp. It supports advanced features including tool calling and follows a specific chat template for multi-turn conversations. The model is intended for developers who want to run powerful language models offline.
- Qwen3.5 2Bhuggingface.coQwen3.5-2B-AWQ-4bit is a 4-bit AWQ quantized version of Alibaba's Qwen 3.5 2B model. It supports multimodal inputs including images and video in addition to text, along with tool calling capabilities. The model is distributed on Hugging Face for efficient local or edge deployment.
- Qwen3.5 35B A3Bhuggingface.coQwen3.5-35B-A3B is an open-source large language model supporting both text and multimodal inputs. It is designed for advanced AI applications, including chatbots and multimodal assistants, and is suitable for developers and researchers in AI.
- Qwen3.6 35B A3Bhuggingface.coThis repository provides GGUF quantized files for the Qwen3.6-35B model using an A3B (likely MoE or distilled) architecture. It is optimized for local inference with tools such as LM Studio. The model includes advanced tool-calling capabilities and a sophisticated chat template for complex interactions.
- Qwen2.5 32B Instructhuggingface.coQwen2.5-32B-Instruct-GPTQ-Int4 is a quantized variant of the Qwen2.5 32B instruction-tuned language model hosted on Hugging Face. It is provided as a GPTQ-Int4 model file intended for inference on compatible hardware. The model follows a system prompt that identifies it as Qwen created by Alibaba Cloud and positions it as a helpful assistant. The page supplies a chat template that defines how the model processes messages. When a system message is present it uses that content; otherwise it defaults to stating that the model is Qwen created by Alibaba Cloud and is a helpful assistant. The template also includes explicit support for tool use. It instructs the model that it may call one or more functions to assist with a user query, supplies function signatures inside XML-style tools tags, and requires each function call to be returned as a JSON object wrapped in tool_call XML tags. This structure enables the model to handle tool calling and function calling formats during generation. The model is distributed through the Hugging Face repository at Qwen/Qwen2.5-32B-Instruct-GPTQ-Int4. It belongs to the class of foundation models made available for download and local or hosted inference.
- Qwen3.5 122B A10Bhuggingface.coThis is a quantized (NVFP4) variant of the Qwen3.5-122B model optimized by NVIDIA for efficient inference. It supports advanced features including tool calling, multimodal inputs, and is designed to run on NVIDIA GPUs with significantly lower memory requirements than the original model. The model is distributed on Hugging Face and can be used with standard Transformers pipelines or custom inference servers.
- Qwen3 4Bhuggingface.coA 4 billion parameter version of the Qwen3 large language model in FP8 precision. It supports instruction following, tool calling, and general text generation. The model is hosted on Hugging Face and can be used locally or via inference providers.
- Qwen3.5 122B A10B Int4 AutoRoundhuggingface.coThis repository hosts an Intel-optimized, AutoRound-quantized version of the Qwen3.5-122B-A10B model in int4 precision. It enables running this large multimodal model on more accessible hardware while maintaining good performance. The model is suitable for advanced AI research and local inference applications.
- Qwen3.5 9Bhuggingface.coQwen3.5-9B-AWQ-4bit is a community-quantized 4-bit AWQ version of the Qwen3.5-9B model. It supports text generation, vision understanding, and tool calling while significantly reducing memory requirements. The model is compatible with Hugging Face Transformers and other inference engines, making advanced multimodal capabilities accessible on standard hardware.
- Qwen3.6 27Bhuggingface.coThis is a 4-bit AWQ quantized version of the Qwen 3.6B model optimized for lower memory usage and faster inference on consumer hardware. It supports multimodal inputs including images and video and includes a chat template for conversational use. The model is distributed on Hugging Face and can be loaded with Transformers or compatible inference engines.
- Qwen 3.6 35B A3B VRAP 4 Bit AWQ 21.2GBhuggingface.coA heavily quantized (4-bit AWQ) version of a Qwen 3.6 model with 35B+3B parameters and vision capabilities (VRAP). The 21.2GB model supports both text and image inputs. It is designed for local inference using GGUF-compatible tools or optimized runtimes.
- Qwen3.6 35B A3Bhuggingface.coQwen/Qwen3.6-35B-A3B-FP8 is an open-source large language model designed for advanced text generation and understanding. It supports instruction following and multilingual capabilities, making it suitable for developers and researchers building AI-powered solutions.
- Qwen3 30B A3Bhuggingface.coQwen3-30B-A3B-FP8 is an FP8 quantized version of Alibaba's Qwen3 series Mixture-of-Experts model. It offers strong performance on reasoning and tool-use tasks while significantly reducing memory requirements compared to full-precision versions. The model supports advanced features such as function calling and multi-step tool use and is distributed for use with Transformers and compatible inference engines.
- Qwen3.5 9B Quantized.w8a8huggingface.coThis repository contains a w8a8 quantized version of the Qwen3.5-9B model from RedHatAI. It supports both text and vision inputs and is optimized for lower memory usage and faster inference. The model includes vision tower components and is compatible with the Transformers library.
- Qwen3.5 122B A10B Int4 Fp8 Hybridhuggingface.coThis is a hybrid int4/fp8 quantized version of the Qwen3.5-122B-A10B model with full support for image and video inputs. It features sophisticated content rendering for multimodal messages and vision counting. The model balances performance and efficiency for large-scale multimodal applications.
- Qwen3.6 27Bhuggingface.coThis is a community-created 8-bit GPTQ quantized version of the Qwen3.6-27B model. It allows the large model to run on consumer GPUs with reduced memory requirements while preserving most of the original performance. The model supports the standard Qwen chat template and is suitable for local deployment using Hugging Face transformers or compatible inference engines.
- Qwen3 4Bhuggingface.coThis is the GGUF quantized format of Alibaba's Qwen3-4B large language model. It enables efficient local inference using tools like llama.cpp or LM Studio. The model includes support for tool calling and follows a specific chat template for structured interactions.
- Qwen3.5 122B A10Bhuggingface.coQwen3.5-122B-A10B-NVFP4 is a highly quantized version of the large Qwen 3.5 model (122B parameters). It supports vision inputs alongside text and includes advanced tool-calling features. The model is distributed on Hugging Face for use in high-performance inference environments.
- Qwen3.6 35B A3Bhuggingface.coQwen3.6-35B-A3B-NVFP4 is a quantized foundation model from the Qwen series, optimized by Unsloth for NVFP4 precision. It supports text and vision inputs and is designed for high-performance local inference. The model is distributed on Hugging Face and targets developers building on-device or cost-efficient multimodal applications.
- Qwen3.6 35B A3B APEX MTPhuggingface.comudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF provides GGUF quantized weights for a Qwen3.6 model incorporating A3B APEX MTP (Mixture-of-Experts or similar advanced technique) optimizations. It is designed for efficient local execution using tools such as llama.cpp. The model includes complex tokenizer configurations for multimodal or advanced prompting capabilities.
- Qwen3.6 35B A3Bhuggingface.coThis is a community-quantized 8-bit version of Qwen3.6-35B (with A3B MoE architecture) prepared for the MLX framework on Apple devices. It enables high-performance local inference on Macs with reduced memory footprint while retaining strong reasoning capabilities. The model uses standard MLX conversion and loading patterns.
- Qwen3.6 35B A3B PrismaQuant 4.75bit Vllmhuggingface.coThis is a 4.75-bit PrismaQuant version of a Qwen3.6 35B-A3B model optimized for use with the vLLM inference engine. It supports multimodal inputs including vision and includes custom chat templates for tool use. The model is distributed on Hugging Face for efficient local or server-based deployment.
- Qwen3.6 35B A3Bhuggingface.coThis is a community-quantized version of the Qwen3.6-35B model using NVFP4 precision. It supports both text and vision inputs and is optimized for reduced memory usage and faster inference on compatible hardware. The model is intended for local deployment and experimentation.
- Qwen Qwen3.6 35B A3Bhuggingface.coThis is a GGUF-quantized version of the Qwen3.6-35B-A3B model, optimized for efficient local inference using tools like llama.cpp. It supports multimodal inputs and is designed for developers who want to run powerful language models on standard hardware without relying on cloud APIs.
- Qwen3 32Bhuggingface.coQwen3-32B-FP8 is an FP8-quantized version of Alibaba's Qwen3 32B large language model. It supports advanced capabilities including tool calling, long-context understanding, and multilingual performance while significantly reducing VRAM usage compared to the original. The model is designed for local and self-hosted inference by developers building AI applications.
- Qwen2.5 32B Instructhuggingface.coQwen2.5-32B-Instruct-GPTQ-Int8 is an INT8-quantized version of Alibaba's Qwen2.5 32B Instruct model. It supports advanced features including tool calling and follows a detailed chat template. The model is designed for efficient inference while retaining strong reasoning and instruction-following capabilities.
- Qwen3.5 9Bhuggingface.coQwen3.5-9B-NVFP4 is a community-quantized version of the Qwen 3.5 9B large language model, optimized for efficient local and on-device inference. It provides GGUF files suitable for tools like llama.cpp and Ollama, supporting text generation, multimodal inputs, and tool-calling capabilities. Primarily used by developers and AI researchers who self-host open-weight models.
- Qwen3 8Bhuggingface.coQwen3-8B-GGUF is the GGUF-quantized format of Alibaba's Qwen3-8B model. This version enables efficient local inference on consumer hardware using tools like llama.cpp or Ollama. It retains strong language understanding and generation capabilities while supporting tool calling and structured prompting formats.
- Qwen3.6 35B A3B CompQuanthuggingface.coQwen3.6-35B-A3B-CompQuant-MLX-3bit is an open-source, quantized large language model optimized for local inference. It features 3-bit compression and MLX compatibility, making it suitable for researchers and developers needing efficient LLMs.
- Qwen3.6 35B A3Bhuggingface.coThis repository provides GGUF quantized files for the Qwen3.6-35B-A3B model, enabling efficient CPU and GPU inference using tools like llama.cpp. The model supports multimodal inputs including vision and offers strong reasoning capabilities. It is designed for users who want to run powerful language models locally without relying on cloud APIs.
Ranked by how close each one sits to Qwen3.5 35B A3B in the index, not by popularity. Back to Qwen3.5 35B A3B →