UM  > 科技學院
Robust log-energy estimation and its dynamic change enhancement for in-car speech recognition
Li W.3; Wang L.1; Zhou Y.2; Bourlard H.4; Liao Q.3
2013-05-22
Source PublicationIEEE Transactions on Audio, Speech and Language Processing
ISSN15587916
Volume21Issue:8Pages:1689-1698
Abstract

The log-energy parameter, typically derived from a full-band spectrum, is a critical feature commonly used in automatic speech recognition (ASR) systems. However, log-energy is difficult to estimate reliably in the presence of background noise. In this paper, we theoretically show that background noise affects the trajectories of not only the 'conventional' log-energy, but also its delta parameters. This results in a poor estimation of the actual log-energy and its delta parameters, which no longer describe the speech signal. We thus propose a new method to estimate log-energy from a sub-band spectrum, followed by dynamic change enhancement and mean smoothing. We demonstrate the effectiveness of the proposed log-energy estimation and its post-processing steps through speech recognition experiments conducted on the in-car CENSREC-2 database. The proposed log-energy (together with its corresponding delta parameters) yields an average improvement of 32.8% compared with the baseline front-ends. Moreover, it is also shown that further improvement can be achieved by incorporating the new Mel-Frequency Cepstral Coefficients (MFCCs) obtained by non-linear spectral contrast stretching. © 2006-2012 IEEE.

KeywordDynamic Change Enhancement In-car Speech Recognition Log-energy Mel-filterbank (Mfb) Mel-frequency Cepstral Coefficients (Mfccs)
DOIhttp://doi.org/10.1109/TASL.2013.2260151
URLView the original
Indexed BySCI
Language英语
WOS Research AreaAcoustics ; Engineering
WOS SubjectAcoustics ; Engineering, Electrical & Electronic
WOS IDWOS:000319020800004
Fulltext Access
Citation statistics
Cited Times [WOS]:3   [WOS Record]     [Related Records in WOS]
Document TypeJournal article
CollectionFaculty of Science and Technology
DEPARTMENT OF COMPUTER AND INFORMATION SCIENCE
Corresponding AuthorZhou Y.
Affiliation1.Nagaoka University of Technology
2.Universidade de Macau
3.Tsinghua University
4.Swiss Federal Institute of Technology, Lausanne
Recommended Citation
GB/T 7714
Li W.,Wang L.,Zhou Y.,et al. Robust log-energy estimation and its dynamic change enhancement for in-car speech recognition[J]. IEEE Transactions on Audio, Speech and Language Processing,2013,21(8):1689-1698.
APA Li W.,Wang L.,Zhou Y.,Bourlard H.,&Liao Q..(2013).Robust log-energy estimation and its dynamic change enhancement for in-car speech recognition.IEEE Transactions on Audio, Speech and Language Processing,21(8),1689-1698.
MLA Li W.,et al."Robust log-energy estimation and its dynamic change enhancement for in-car speech recognition".IEEE Transactions on Audio, Speech and Language Processing 21.8(2013):1689-1698.
Files in This Item:
There are no files associated with this item.
Related Services
Recommend this item
Bookmark
Usage statistics
Export to Endnote
Google Scholar
Similar articles in Google Scholar
[Li W.]'s Articles
[Wang L.]'s Articles
[Zhou Y.]'s Articles
Baidu academic
Similar articles in Baidu academic
[Li W.]'s Articles
[Wang L.]'s Articles
[Zhou Y.]'s Articles
Bing Scholar
Similar articles in Bing Scholar
[Li W.]'s Articles
[Wang L.]'s Articles
[Zhou Y.]'s Articles
Terms of Use
No data!
Social Bookmark/Share
All comments (0)
No comment.
 

Items in the repository are protected by copyright, with all rights reserved, unless otherwise indicated.