Overview of InfiniteTalk
mainInfiniteTalk is an audio-driven video generation model designed for sparse-frame video dubbing. It supports two primary modes of operation:
- Video-to-Video: Given an input video and an audio track, it synthesizes a new video with accurate lip synchronization while aligning head movements, body posture, and facial expressions with the audio.
- Image-to-Video: Uses an image and an audio track as input to generate a talking video.
Key capabilities include infinite-length video generation, consistent identity preservation, and improved stability in body/hand movements compared to previous methods like MultiTalk.