Stable Audio 3 Medium is a Stability AI text-to-audio model for generating music and sound effects. The model is distributed through Hugging Face with Safetensors files and can be used locally or through supported inference providers by developers and audio creators.
In the Text to speech space, Stable Audio 3 Medium takes a focused approach. It focuses on generating music and sound effects from text prompts without creating them manually. Stable Audio 3 Medium is an open-source project aimed at audio developers, music producers, and AI researchers. Stable Audio 3 Medium is open source under the MIT license. It ships for the web and API, and it can be self-hosted.
Stability AI builds and maintains Stable Audio 3 Medium, and it first shipped in 2026. Development happens publicly on GitHub with 713 stars and 143 commits in the last 90 days. Key capabilities include text-to-audio, music generation, and sound effects. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do