Whisper To Stable Diffusion is a web application that transcribes spoken prompts using speech-to-text AI and then generates images from the transcribed text using image generation models. It is designed for AI enthusiasts and creators who want to explore multimodal AI workflows.
In the Other AI space, Whisper To Stable Diffusion takes a focused approach. Enabling users to generate images from spoken prompts by transcribing speech to text and using AI image models. It is built as a consumer product for AI enthusiasts and creators. It is available for free. It ships for the web, and it can be self-hosted.
fffiloni builds and maintains Whisper To Stable Diffusion, and it first shipped in 2023. Among its 4 catalogued features are speech-to-text, image generation, and prompt transcription.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do