LLMKube is a Kubernetes-native operator that simplifies deploying and managing large language model inference services on consumer or enterprise GPUs. It supports multiple backends including vLLM, llama.cpp, and TGI across NVIDIA, Apple Silicon, and AMD hardware. It includes a CLI for rapid deployment and an agentic harness called Foreman that enables local models to autonomously contribute code via pull requests.
In the AI & ML space, LLMKube takes a focused approach. It focuses on running production-grade LLM inference on local or self-managed hardware with Kubernetes orchestration. LLMKube is an open-source project aimed at developers. The project is open source (Apache-2.0). It runs on the command line, and it can be self-hosted.
It is developed by defilantech, and the product first shipped in 2025. The project is developed in the open on GitHub with 175 stars and 503 commits in the last 90 days. Among its 6 catalogued features are Kubernetes Operator, vLLM Support, and llama.cpp Support.
Latest indexed changes and source events
defilantech/LLMKube verified by the PulseGate indexer
Other apps tracked under the same category.