LatentSync-1.6 is a machine learning model developed by ByteDance for audio-driven video lip synchronization. It generates realistic mouth movements, teeth, and facial expressions from an input audio track and reference video. The model was trained on higher resolution (512x512) videos to address blurriness issues in prior versions and is available for use through Hugging Face with various inference libraries.
In the Other AI space, LatentSync takes a focused approach. It focuses on creating realistic lip movements and facial expressions in video that match given audio input. It is built as an open-source project for AI researchers and video developers. LatentSync is open source under the Apache-2.0 license. It ships for the web and API.
Behind LatentSync is ByteDance, and it first shipped in 2024. Development happens publicly on GitHub with 5.9k stars. Key capabilities include lip sync generation, high-resolution training, and video editing.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do