3D Lip-Synch Generation with Data-Faithful Machine Learning

Kim, Ig-Jae; Ko, Hyeong-Seok

3D Lip-Synch Generation with Data-Faithful Machine Learning

Date

2007

Authors

Kim, Ig-Jae
Ko, Hyeong-Seok

Publisher

The Eurographics Association and Blackwell Publishing Ltd

Abstract

This paper proposes a new technique for generating three-dimensional speech animation. The proposed technique takes advantage of both data-driven and machine learning approaches. It seeks to utilize the most relevant part of the captured utterances for the synthesis of input phoneme sequences. If highly relevant data are missing or lacking, then it utilizes less relevant (but more abundant) data and relies more heavily on machine learning for the lip-synch generation. This hybrid approach produces results that are more faithful to real data than conventional machine learning approaches, while being better able to handle incompleteness or redundancy in the database than conventional data-driven approaches. Experimental results, obtained by applying the proposed technique to the utterance of various words and phrases, show that (1) the proposed technique generates lip-synchs of different qualities depending on the availability of the data, and (2) the new technique produces more realistic results than conventional machine learning approaches.

        @article{10.1111:j.1467-8659.2007.01051.x
,
journal = {Computer Graphics Forum},
title = {{3D Lip-Synch Generation with Data-Faithful Machine Learning
}},
author = {Kim, Ig-Jae and 
Ko, Hyeong-Seok
},
year = {2007
},
publisher = {The Eurographics Association and Blackwell Publishing Ltd
},
ISSN = {1467-8659
},
DOI = {10.1111/j.1467-8659.2007.01051.x
}
}

URI

https://doi.org/10.1111/j.1467-8659.2007.01051.x

Collections

26-Issue 3

Full item page