Warp is a dependency-free, embeddable C inference engine for running very large language models locally. It streams activated model weights directly from NVMe storage, enabling models such as Kimi K3 and GLM-5.3-Flash to run beyond available RAM.
In the Inference & model serving space, Warp takes a focused approach. It focuses on running large language models locally when their activated weights exceed available system RAM. Warp is an open-source project aimed at developers building local AI inference applications. The project is open source (Open Source). Warp is available on the web, embeddable surfaces, the command line, macOS, and Linux, and it can be self-hosted.
sqliteai builds and maintains Warp. Key capabilities include NVMe weight streaming, C inference engine, and dependency-free runtime.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match