JANGQ improves quantization of large language models by classifying tensors into sensitivity tiers and applying different bit widths: higher precision (5-8 bits) for attention layers and aggressive compression (2-4 bits) for MLP layers. This results in smaller models with superior performance compared to uniform MLX baselines, as demonstrated in curated benchmarks showing large MMLU gains with reduced memory. The open-source project includes runtime, profiles, and detailed results for models like MiniMax, Qwen, and Mistral.
JANGQ is an AI & ML project. Standard uniform quantization degrading performance in smaller LLMs by treating all layers with equal precision regardless of sensitivity. JANGQ is an open-source project aimed at machine learning engineers and researchers. The project is open source (Open Source). JANGQ is available on the command line.
Behind JANGQ is JANG, and it first shipped in 2026. The project is developed in the open on GitHub with 218 stars and 149 commits in the last 90 days. Among its 5 catalogued features are Variable Bit Quantization, Layer Sensitivity Analysis, and MLX Optimization. The interface is available in 5 languages, including English, Spanish, and Japanese.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do