forewarm is an open-source tool for predictive autoscaling of LLM inference workloads based on token demand forecasts. It is intended for infrastructure teams operating vLLM and Kubernetes environments, with integrations or deployment patterns involving KEDA.
In the Inference & model serving space, forewarm takes a focused approach. It focuses on preventing inefficient or delayed scaling of LLM inference workloads under changing token demand. It is built as an open-source project for ML infrastructure and platform engineers. forewarm is open source under the Apache-2.0 license. It ships for the command line, and it can be self-hosted.
Behind forewarm is Alexey Kazantsev, and it first shipped in 2026. The project is developed in the open on GitHub with 1 commit in the last 90 days. Among its 6 catalogued features are token forecasting, predictive autoscaling, and LLM inference scaling. forewarm is currently in beta.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match