This is a quantized (NVFP4) variant of the Qwen3.5-122B model optimized by NVIDIA for efficient inference. It supports advanced features including tool calling, multimodal inputs, and is designed to run on NVIDIA GPUs with significantly lower memory requirements than the original model. The model is distributed on Hugging Face and can be used with standard Transformers pipelines or custom inference servers.
Qwen3.5 122B A10B is a Foundation models & chat project. It focuses on deploying and running a large 122B-parameter language model efficiently on NVIDIA hardware with reduced memory footprint. It is built as an open-source project for AI developers and researchers. Qwen3.5 122B A10B is open source under the Apache-2.0 license. It runs on the web and API.
Behind Qwen3.5 122B A10B is NVIDIA, and it first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 354 commits in the last 90 days. The category is crowded — PulseGate's index counts 23 comparable projects. Key capabilities include Text Generation, Quantized Weights, and Tool Calling.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do