inkling-GGUF provides pre-quantized weights of the Inkling model in GGUF format for efficient local inference. It is designed for use with llama.cpp, Ollama and other GGUF-compatible runtimes. The model enables developers to run capable language models on CPUs or modest GPUs with reduced memory requirements.
Inkling is a Foundation models & chat project. It focuses on running large language models efficiently on consumer hardware without cloud dependency. Inkling is an open-source project aimed at developers. The project is open source (Apache-2.0). It ships for the web, the command line, and API.
It is developed by Unsloth, and it first shipped in 2023. Development happens publicly on GitHub with 68.8k stars and 1.2k commits in the last 90 days. Among its 3 catalogued features are quantized weights, GGUF format, and local inference.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do