This repository hosts a 4.75-bit quantized version of a 35B-parameter (with 3.6B active) multimodal model optimized for vLLM on NVIDIA Blackwell GPUs. It includes custom chat templates supporting vision and video inputs. The model is intended for developers seeking efficient local or server-based multimodal inference.
75bits is a Foundation models & chat product. It focuses on delivering a compact, high-performance multimodal model that runs efficiently on NVIDIA Blackwell hardware using vLLM. 75bits is an open-source project aimed at AI inference engineers. The project is open source (Open Source). The product ships for the web, the command line, and API.
Behind 75bits is cyburn, and the product first shipped in 2026. The project is developed in the open on GitHub with 94 stars and 616 commits in the last 90 days. Among its 4 catalogued features are quantized weights, vLLM optimized, and multimodal.
Latest indexed changes and source events
cyburn/Qwopus3.6-35B-A3B-v1-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm-4.75bits verified by the PulseGate indexer
Other apps tracked under the same category.