This is an AWQ 4-bit quantized version of the TinyLlama 1.1B Chat model, created by TheBloke. It provides a small, efficient language model suitable for local deployment on CPUs or low-end GPUs. The model uses a chat template and is ideal for experimentation, edge deployment, or resource-constrained environments.
In the Foundation models & chat space, TinyLlama 1.1B Chat takes a focused approach. It focuses on running capable chat models on devices with limited memory and compute. TinyLlama 1.1B Chat is an open-source project aimed at developers. The project is open source (Apache-2.0). It runs on the web, the command line, and API.
TheBloke builds and maintains TinyLlama 1.1B Chat, and the product first shipped in 2023. The project is developed in the open on GitHub with 86.6k stars and 3k commits in the last 90 days. Among its 3 catalogued features are 4-bit AWQ quantization, chat format, and low-resource inference.
Latest indexed changes and source events
TheBloke/TinyLlama-1.1B-Chat-v0.3-AWQ verified by the PulseGate indexer
Other apps tracked under the same category.