This is a quantized (NVFP4) variant of the Qwen3.5-122B model optimized by NVIDIA for efficient inference. It supports advanced features including tool calling, multimodal inputs, and is designed to run on NVIDIA GPUs with significantly lower memory requirements than the original model. The model is distributed on Hugging Face and can be used with standard Transformers pipelines or custom inference servers.
Qwen3.5 122B A10B is a Foundation models & chat product. It focuses on deploying and running a large 122B-parameter language model efficiently on NVIDIA hardware with reduced memory footprint. Qwen3.5 122B A10B is an open-source project aimed at AI developers and researchers. The project is open source (Apache-2.0). The product ships for the web and API.
NVIDIA builds and maintains Qwen3.5 122B A10B, and the product first shipped in 2024. The project is developed in the open on GitHub with 3.3k stars and 354 commits in the last 90 days. It competes in a saturated segment with 23 similar apps in PulseGate's index. Among its 4 catalogued features are Text Generation, Quantized Weights, and Tool Calling.
Latest indexed changes and source events
nvidia/Qwen3.5-122B-A10B-NVFP4 verified by the PulseGate indexer
Other apps tracked under the same category.