Speech gesture generation from the trimodal context of text, audio, and speaker identity

Explore this paper's citation graph

Summary

This paper presents an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures that are human-like and that match with speech content and rhythm.

Type
article
Published
2020-09-04
Cited by
400
References
70
Access
Open access

Keywords

Gesture, Computer science, Context (archaeology), Speech recognition, Identity (music)

References

Cited by

Related papers