LlamaRack
PulseGate's liveness check found it on 6 Oct 2026; it has been in the index since 5 Sep 2026. How this is checked
LlamaRack is a self-hosted management platform for llama.cpp runtimes. It provides a web UI, multi-model orchestration, GPU-aware scheduling, monitoring, runtime management, and OpenAI-compatible APIs for developers operating local language models.
Inferred · not functionally tested
Overview
6 featuresPurpose: Managing, scheduling, monitoring, and exposing multiple self-hosted llama.cpp model runtimes.
Inferred · not functionally tested
Audience: developers and infrastructure teams running local LLMs
Inferred · not functionally tested
Functions: agents, monitoring
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: unknown · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Open Source · platforms: API, WEB · deployment: browser, self_hosted, api_only
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: github.com. These links do not verify the individual claims.
LlamaRack sits in PulseGate's Inference & model serving category. Inferred · not functionally tested: Managing, scheduling, monitoring, and exposing multiple self-hosted llama.cpp model runtimes. Inferred · not functionally tested: It is built as an open-source project for developers and infrastructure teams running local LLMs. Basis unknown · not verified: The project is open source (Open Source). Basis unknown · not verified: LlamaRack is available on the web and API, and it can be self-hosted.
It is developed by brantje. Inferred · not functionally tested: Key capabilities include modern web UI, multi-model orchestration, and GPU scheduling. Inferred · not functionally tested: Catalogued interfaces include a public API.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Modern web UI
- Multi-model orchestration
- GPU scheduling
- OpenAI-compatible API
- Runtime management
- Monitoring
Topics: Inferred · not functionally tested
Built with & integrations
- local_oss
- bllama in the HTML
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed5 Sep · 23:42 UTCAnyone using llama.cpp willing to test LlamaRack? seen via Hacker News firehose (Algolia)Source: Hacker News firehose (Algolia) · Open
Frequently asked questions about LlamaRack
- What does LlamaRack do?
- Inferred · not functionally tested: Managing, scheduling, monitoring, and exposing multiple self-hosted llama.cpp model runtimes. It is catalogued under Inference & model serving on PulseGate.
- Who is LlamaRack for?
- Inferred · not functionally tested: LlamaRack is an open-source project built for developers and infrastructure teams running local LLMs.
- Does LlamaRack have a free plan?
- Basis unknown · not verified: Yes — LlamaRack is open source under the Open Source license and free to use.
- What platforms does LlamaRack run on?
- Basis unknown · not verified: LlamaRack runs on the web and API. It can also be self-hosted.
- Is LlamaRack still active?
- PulseGate's liveness check found it on 6 Oct 2026.
- What are alternatives to LlamaRack?
- Similar projects tracked by PulseGate include LlamaStudio, The Llama.cpp Fork That Enables Qwen 3.8 27B Large Contexts for 16GB VRAM GPU, and llama_cpp.rb.LlamaStudioThe Llama.cpp Fork That Enables Qwen 3.8 27B Large Contexts for 16GB VRAM GPUllama_cpp.rb
- Who makes LlamaRack?
- LlamaRack is developed by brantje.
- Is LlamaRack open source?
- Basis unknown · not verified: Yes — LlamaRack is open source under the Open Source license.
Similar projects
Closest matches by what these projects do