MOVA-360p is an open multimodal diffusion model hosted on Hugging Face for image-to-video generation. It accepts an input image together with an optional text prompt and produces a short video clip that matches both the visual content and the described motion or scene.
The model supports several related tasks listed on its repository page: image-text-to-video, image-to-audio-video, and image-text-to-audio-video. It is distributed with Safetensors weights and carries an Apache-2.0 license. Integration with the Diffusers library is provided through a standard pipeline that loads the model in bfloat16 precision, runs inference on CUDA, and returns video frames ready for export.
Developers can install the required packages with pip and instantiate the pipeline using a few lines of Python code that loads an image from a URL or local path and combines it with a descriptive prompt. The repository also references an associated arXiv paper for technical details. No pricing information appears because the model is offered as a free, open-source download.
MOVA 360p is a Video generation product. It focuses on converting images and text into synchronized video and audio outputs using a single model. It is built as an open-source project for developers. MOVA 360p is open source under the Apache-2.0 license. The product ships for the web, the command line, and API.
Behind MOVA 360p is OpenMOSS-Team, and the product first shipped in 2026. Development happens publicly on GitHub with 1.1k stars and 5 commits in the last 90 days. Key capabilities include image-to-Video, image-to-Audio, and Diffusers Pipeline. It exposes integrations via a public API.
Latest indexed changes and source events
OpenMOSS-Team/MOVA-360p verified by the PulseGate indexer
Other apps tracked under the same category.