MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
Yue Zhang
*
, Zhizhou Zhong
*
, Minhao Liu
*
, Zhaokang Chen, Bin Wu
†
, Yubin Zeng, Chao Zhang, Yingjie He, Junxin Huang, Wenjiang Zhou
(
*
Equal Contribution,
†
Corresponding Author, benbinwu@tencent.com) Lyra Lab, Tencent Music Entertainment
[Github Repo]
[Huggingface]
[Technical report]
Drving Audio
Drop Audio Here
- or -
Click to Upload
Reference Image or Video
Drop File Here
- or -
Click to Upload
BBox_shift value, px
Extra Margin
↺
0
40
Parsing Mode
jaw
raw
Left Cheek Width
↺
20
160
Right Cheek Width
↺
20
160
'left_cheek_width' and 'right_cheek_width' parameters determine the range of left and right cheeks editing when parsing model is 'jaw'. The 'extra_margin' parameter determines the movement range of the jaw. Users can freely adjust these three parameters to obtain better inpainting results.
1. Test Inpainting
2. Generate
Test Inpainting Result (First Frame)
Parameter Information
Video