This is a GPTQ 4-bit quantized version of TinyLlama-1.1B-Chat-v1.0 by TheBloke. It provides an efficient way to run the 1.1 billion parameter chat model locally with significantly reduced memory requirements while maintaining good performance.
TinyLlama 1.1B Chat sits in PulseGate's Quantised & converted weights category. It focuses on running a capable small language model for chat on resource-constrained devices through 4-bit GPTQ quantization. TinyLlama 1.1B Chat is an open-source project aimed at Local LLM users and developers. TinyLlama 1.1B Chat is open source under the AGPL-3.0 license. It runs on the web, the command line, and API.
It is developed by TheBloke, and it first shipped in 2022. Development happens publicly on GitHub with 47.5k stars and 141 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 5 similar projects. Key capabilities include GPTQ quantization, chat format support, and small footprint.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do