Dithering techniques in automatic recognition of speech corrupted by MP3 compression: Analysis, solutions and experiments

Michal Borsky*, Petr Mizera, Petr Pollak, Jan Nouza

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

3 Citations (Scopus)

Abstract

A large portion of the audio files distributed over the Internet or those stored in personal and corporate media archives are in a compressed form. There exist several compression techniques and algorithms but it is the MPEG Layer-3 (known as MP3) that has achieved a really wide popularity in general audio coding, and in speech, too. However, the algorithm is lossy in nature and introduces distortion into spectral and temporal characteristics of a signal. In this paper we study its impact on automatic speech recognition (ASR). We show that with decreasing MP3 bitrates the major source of ASR performance degradation is deep spectral valleys (i.e. bins with almost zero energy) caused by the masking effect of the MP3 algorithm. We demonstrate that these unnatural gaps in spectrum can be effectively compensated by adding a certain amount of noise to the distorted signal. We provide theoretical background for this approach where we show that the added noise affects mainly the spectral valleys. They are filled by the noise while the spectral bins with speech remain almost unchanged. This helps to restore a more natural shape of log spectrum and cepstrum, and consequently has a positive impact on ASR performance. In our previous work, we have proposed two types of the signal dithering (noise addition) technique, one applied globally, the other in a more selective way. In this paper, we offer a more detailed insight into their performance. We provide results from many experiments where we test them in various scenarios, using a large vocabulary continuous speech recognition (LVCSR) system, acoustic models based on gaussian-mixture model (GMM) as well as on deep-neural network (DNN), and multiple speech databases in three languages (Czech, English and German). Our results prove that both the proposed techniques, and the selective dithering method, in particular, yield consistent compensation of the negative impact of the MP3 compressed speech on ASR performance.

Original languageEnglish
Pages (from-to)75-84
Number of pages10
JournalSpeech Communication
Volume86
DOIs
Publication statusPublished - 1 Feb 2017

Bibliographical note

Funding Information:
The research described in the paper was supported by CTU Grant SGS14/191/OHK3/3T/13 “Advanced Algorithms of Digital Signal Processing and their Applications” and by Technology Agency of the Czech Republic in project no. TA04010199 called “MultiLinMedia”.

Publisher Copyright:
© 2016 Elsevier B.V.

Other keywords

  • DNN–HMM
  • GMM–HMM
  • MP3 compression
  • Spectrally selective dithering
  • Uniform dithering

Fingerprint

Dive into the research topics of 'Dithering techniques in automatic recognition of speech corrupted by MP3 compression: Analysis, solutions and experiments'. Together they form a unique fingerprint.

Cite this