pyinferencemanager is a Python library that acts as an intelligent orchestrator for AI inference workloads. It provides multi-provider routing across services like OpenAI, Anthropic, and local models (Ollama), along with semantic caching, dynamic routing, cost optimization, budget enforcement, and hardware-aware scheduling. Designed for developers building LLM-powered applications, it simplifies managing distributed inference with built-in tracing and optimization features.
pyinferencemanager is an AI & ML project. It focuses on managing and routing AI inference workloads across multiple LLM providers while controlling costs, enforcing budgets, and optimizing performance. pyinferencemanager is an open-source project aimed at developers. The project is open source (MIT). It runs on the command line and API.
Behind pyinferencemanager is Mullassery, and it first shipped in 2026. The project is developed in the open on GitHub with 15 commits in the last 90 days. Among its 7 catalogued features are multi-provider routing, semantic caching, and cost optimization. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match