Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › глава/раздел › Рецензирование
Microphone array post-filter in frequency domain for speech recognition using short-time log-spectral amplitude estimator and spectral harmonic/noise classifier. / Salishev, Sergey; Klotchkov, Ilya; Barabanov, Andrey.
Speech and Computer (SPECOM 2017). Springer Nature, 2017. стр. 525-534 (Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); Том 10458 LNAI).Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › глава/раздел › Рецензирование
}
TY - CHAP
T1 - Microphone array post-filter in frequency domain for speech recognition using short-time log-spectral amplitude estimator and spectral harmonic/noise classifier
AU - Salishev, Sergey
AU - Klotchkov, Ilya
AU - Barabanov, Andrey
PY - 2017/1/1
Y1 - 2017/1/1
N2 - We propose a novel computationally efficient real-time microphone array speech enhancement postfilter with a small delay that takes into account features of speech signal and recognition algorithms. The algorithm is efficient for small microphone arrays. The filter is based on applying a binary classification model to the Log Short-Term Spectral Amplitude (Log-STSA). The proposed algorithm allows substantial improvement of recognition accuracy with minor increase in complexity compared to Wiener post-filter and lower complexity compared to existing voice model based approaches. Objective tests using dual microphone array, ETSI binaural noise database, TIDIGITS database, and CMU Sphinx 4 speech recognizer demonstrate overall 41% Error Rate reduction for SNR from 15 dB to 0 dB. Subjective evaluation also demonstrates substantial noise reduction and intelligibility improvement without musical noise artifacts common for Wiener and Spectral Subtraction based methods. Testing with SiSEC10 four microphone linear equispaced array database shows that recognition accuracy is improved with increased base and/or number of microphones in array.
AB - We propose a novel computationally efficient real-time microphone array speech enhancement postfilter with a small delay that takes into account features of speech signal and recognition algorithms. The algorithm is efficient for small microphone arrays. The filter is based on applying a binary classification model to the Log Short-Term Spectral Amplitude (Log-STSA). The proposed algorithm allows substantial improvement of recognition accuracy with minor increase in complexity compared to Wiener post-filter and lower complexity compared to existing voice model based approaches. Objective tests using dual microphone array, ETSI binaural noise database, TIDIGITS database, and CMU Sphinx 4 speech recognizer demonstrate overall 41% Error Rate reduction for SNR from 15 dB to 0 dB. Subjective evaluation also demonstrates substantial noise reduction and intelligibility improvement without musical noise artifacts common for Wiener and Spectral Subtraction based methods. Testing with SiSEC10 four microphone linear equispaced array database shows that recognition accuracy is improved with increased base and/or number of microphones in array.
KW - Beamforming
KW - Noise reduction
KW - Postfilter
KW - Speech recognition
UR - http://www.scopus.com/inward/record.url?scp=85029481792&partnerID=8YFLogxK
U2 - 10.1007/978-3-319-66429-3_52
DO - 10.1007/978-3-319-66429-3_52
M3 - Chapter
AN - SCOPUS:85029481792
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 525
EP - 534
BT - Speech and Computer (SPECOM 2017)
PB - Springer Nature
T2 - 19th International Conference on Speech and Computer
Y2 - 11 September 2017 through 15 September 2017
ER -
ID: 152227169