Skip to content
AI
Inference & model serving
Inference & model serving inside AI.
Newest through the gate
- yongapypi.orgYonga is a Linux-first local AI model manager for Intel AI PCs.
- freewhirrpypi.orgfreewhirr is an OpenRouter router and OpenAI-compatible proxy for routing LLM requests.
- zhongyitangpypi.orgzhongyitang is a CLI for setting up local LLM inference and an OpenAI-compatible quota gateway.
- SpicyAPIspicyapi.aiSpicyAPI provides hosted access to uncensored image, video, text, and audio AI models through an API and browser studio.
- ScoutWyze Computegithub.comScoutWyze Compute recommends and routes AI workloads to suitable live GPU offers for developers and AI agents.
- AhuraSense Cloudahurasense.comAhuraSense Cloud provides serverless AI inference, GPU compute, and cloud infrastructure for developers and teams.
- mflux-teacachepypi.orgmflux-teacache provides TeaCache step skipping for the mflux command-line diffusion tool.
- TokenRoutergithub.ioTokenRouter is a serving engine for token-level routing between language models.
- GPUYardgpuyard.comGPUYard provides rentable dedicated GPU servers for AI, machine learning, rendering, and other compute workloads.
- GPUniqgpuniq.comGPUniq provides rented GPU infrastructure, API routing, deployment, and monitoring for AI workloads.
- Simplismartsimplismart.aiSimplismart provides scalable AI model inference and deployment for businesses and developers.
- Qubax AIqubax.aiQubax AI provides a unified OpenAI-compatible API for accessing discounted AI models.
- AnimaRenderanimarender.comAnimaRender provides cloud CPU and GPU rendering for 3D projects and popular modeling software.
- Fox Renderfarmfoxrenderfarm.comFox Renderfarm provides cloud-based 3D rendering for animation, VFX, and architectural visualization teams.
- RebusFarmrebusfarm.netRebusFarm provides cloud-based CPU and GPU rendering for 3D images and animations.
- carmen-kernelspypi.orgCarmen-kernels uses AI agents and benchmarking to generate optimized Apple Metal GPU kernels for model inference.
- KryptonLabkryptonlab.idKryptonLab provides a unified AI gateway and API key for accessing models from multiple providers.
- ModelRushmodelrush.aiModelRush provides an OpenAI-compatible API for accessing text, image, video, and voice models.
- BoostRailboostrail.comBoostRail provides a unified API for accessing text, image, and video AI models.
- Cruncrun.aiCrun provides a unified API for accessing video, image, audio, and text AI models.
- model-serving-minefieldpypi.orgmodel-serving-minefield provides read-only diagnostics for model-serving failures.
- InBoost Proxyinboost.proInBoost Proxy is a local inference proxy and reliability gateway for AI coding tools.
- Deepseek V4 1 Flashruninfra.aiDeepSeek V4.1 Flash API on RunInfra: 25% off until Oct 13, 2026, 5:44 AM UTC, $0.10 input and $0.43 output. Standard…
- DEVUP AIdevupai.comDEVUP AI provides an OpenAI-compatible gateway to AI models and on-demand cloud compute for developers.
Watchlist · this browser only, no account
This is the newest 24 of 733. Open Inference & model serving in the live index →
Subscribe to Inference & model serving by RSS — new listings in this category, in your reader, no account.
Elsewhere in AI
Foundation models & chatCoding AI & assistantsImage generationVideo generationVoice, TTS & speechAutonomous agents & workflowsRAG, search & retrievalData science & ML workbenchFine-tuning & trainingLLM eval & observabilityOther AIComputer vision, OCR & document AIWriting & editingAI security & guardrails3D generation