vramscout is an open-source command-line planner for estimating GPU VRAM requirements and maximum context lengths for open-weight LLM deployments. It accounts for model memory and KV-cache constraints and is aimed at developers using inference stacks such as vLLM or SGLang.
vramscout is an Inference & model serving project. It focuses on estimating GPU memory and context limits before deploying open-weight language models. It is built as an open-source project for developers deploying open-weight LLM inference systems. vramscout is open source under the MIT license. It runs on the command line, and it can be self-hosted.
Behind vramscout is S3vryn, and it first shipped in 2026. The project is developed in the open on GitHub with 53 commits in the last 90 days. Among its 8 catalogued features are VRAM estimation, context planning, and KV-cache analysis.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do