Computer Science > Computer Vision and Pattern Recognition

arXiv:2312.15900 (cs)

[Submitted on 26 Dec 2023]

Title:Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

Authors:Zunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li, Xiu Li

Abstract:This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing during inference. To address this problem, we suggest using speech-derived multimodal priors to improve gesture generation. We introduce a novel method that separates priors from speech and employs multimodal priors as constraints for generating gestures. Our approach utilizes a chain-like modeling method to generate facial blendshapes, body movements, and hand gestures sequentially. Specifically, we incorporate rhythm cues derived from facial deformation and stylization prior based on speech emotions, into the process of generating gestures. By incorporating multimodal priors, our method improves the quality of generated gestures and eliminate the need for expensive setup preparation during inference. Extensive experiments and user studies confirm that our proposed approach achieves state-of-the-art performance.

Comments:	AAAI-2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2312.15900 [cs.CV]
	(or arXiv:2312.15900v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2312.15900

Submission history

From: Zunnan Xu [view email]
[v1] Tue, 26 Dec 2023 06:30:14 UTC (1,930 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators