This is a community-quantized W8A8 (weights 8-bit, activations 8-bit) version of the Mistral-Nemo-Instruct-2407 model using dynamic per-token quantization. It is designed for efficient local inference while preserving model quality. The model includes a detailed chat template for instruction following.
In the Foundation models & chat space, Mistral Nemo Instruct 2407 W8A8 Dynamic Per Token takes a focused approach. It focuses on running the Mistral-Nemo instruct model with reduced precision and memory footprint using dynamic per-token quantization. It is built as an open-source project for developers. Mistral Nemo Instruct 2407 W8A8 Dynamic Per Token is open source under the Open Source license. The product ships for the web and API.
noneUsername builds and maintains Mistral Nemo Instruct 2407 W8A8 Dynamic Per Token, and the product first shipped in 2024. Key capabilities include Quantized LLM, 8-bit Weights, and Dynamic Per-Token.
Latest indexed changes and source events
noneUsername/Mistral-Nemo-Instruct-2407-W8A8-Dynamic-Per-Token verified by the PulseGate indexer
Other apps tracked under the same category.