flint-llm is a Python package and CLI tool that simplifies deploying large language models to Kubernetes. It supports popular inference backends including vLLM, Ollama, and Text Generation Inference (TGI). The tool is designed for developers and ML engineers who need to run scalable LLM inference workloads in containerized environments.
flint-llm is a Frameworks & SDKs product. It focuses on deploying and scaling large language models on Kubernetes clusters. It is built as an open-source project for developers. flint-llm is open source under the Apache-2.0 license. It runs on the command line, and it can be self-hosted.
flint-llm first shipped in 2026. Key capabilities include Kubernetes Deployment, LLM Inference, and vLLM Support. It exposes integrations via a public API.
Latest indexed changes and source events
Other apps tracked under the same category.