A study on Korean connected digit recognition and short-term cepstral mean normalization한국어 연속 숫자 음성 인식과 단구간 켑스트럼 평균 정규화에 관한 연구

Cited 0 time in webofscience Cited 0 time in scopus
  • Hit : 654
  • Download : 0
Although many researchers have studied about digit recognition, it is still away from commercial applications in Korea. It is well known that Korean digit recognition is more difficult than English digit recognition, even worse in continuous digits. In this paper, I studied about various techniques to improve the recognition, especially one of the environmental compensation preprocessing methods, called the cepstral mean normalization, with some acoustic-phonetic models. I found that the recognition results varied depending on the windows size for the cepstral mean normalization, and not always the long-term cepstral mean normalization produces the best results. This can be interpreted as if we use the short-term cepstral mean normalization technique with a proper window size for Korean digit recognition, we can get the better results than the conventional cepstral mean normalization. The reason could be the variation of the phone length caused by the short-term cepstral mean normalization, and this variation is believed to improve the recognition rate. Monophone, triphone, whole-word, tri-word, and phonological-rule- considered digit models in Korean pronunciation, are tested in various numbers of states and mixtures. Mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction (PLP) cepstral coefficients are extracted as the feature vectors. Long-term and short-term cepstral mean normalization/ subtraction(CMN/CMS) processing, and relative spectral (RASTA) processing is used for the channel noise compensation. Kalman filtering is applied for additive noise reduction. Linear discriminant analysis (LDA) transformation for the digit recognition is also tested in the end.
Advisors
Hahn, Min-Sooresearcher한민수researcher
Description
한국정보통신대학원대학교 : 공학부,
Publisher
한국정보통신대학원대학교
Issue Date
2002
Identifier
392127/225023 / 020003853
Language
eng
Description

학위논문(석사) - 한국정보통신대학원대학교 : 공학부, 2002, [ xi, 97 p. ]

Keywords

Short-Term Cepstral; Connected Digit Recognition; 인식 시스템; 연속 숫자 음성 인식; ST-CMN

URI
http://hdl.handle.net/10203/54775
Link
http://library.kaist.ac.kr/search/detail/view.do?bibCtrlNo=392127&flag=dissertation
Appears in Collection
School of Engineering-Theses_Master(공학부 석사논문)
Files in This Item
There are no files associated with this item.

qr_code

  • mendeley

    citeulike


rss_1.0 rss_2.0 atom_1.0