TL;DR: Experimental results show that TECCs computed from Gammatone filter bank are more robust in noisy environments than other extracted features, while their performance is practically similar to clean environments.
Abstract: This paper focuses on a robust feature extraction algorithm for automatic classification of pathological and normal voices in noisy environments. The proposed algorithm is based on human auditory processing and the nonlinear Teager-Kaiser energy operator. The robust features which labeled Teager Energy Cepstrum Coefficients (TECCs) are computed in three steps. Firstly, each speech signal frame is passed through a Gammatone or Mel scale triangular filter bank. Then, the absolute value of the Teager energy operator of the short-time spectrum is calculated. Finally, the discrete cosine transform of the log-filtered Teager Energy spectrum is applied. This feature is proposed to identify the pathological voices using a developed neural system of multilayer perceptron (MLP). We evaluate the developed method using mixed voice database composed of recorded voice samples from normophonic or dysphonic speakers. In order to show the robustness of the proposed feature in detection of pathological voices at different White Gaussian noise levels, we compare its performance with results for clean environments. The experimental results show that TECCs computed from Gammatone filter bank are more robust in noisy environments than other extracted features, while their performance is practically similar to clean environments.
TL;DR: This paper focuses on recognising voice corresponding to English alphabets using Mel-frequency cepstral coefficients (MFCCs) and Dynamic Time Warping (DTW) introduced by Sakoe Chiba.
Abstract: This paper focuses on recognising voice corresponding to English alphabets using Mel-frequency cepstral coefficients (MFCCs) and Dynamic Time Warping (DTW) introduced by Sakoe Chiba MFCC are the coefficients that collectively represent the short-term power spectrum of a sound, based on a linear cosine transform of a log power spectrum on a nonlinear mel scale of frequency For recognition samples, voice corresponding to numeric 1 to 5 is taken The endpoint detection, framing, normalization, MFCC and DTW algorithm are then used to process these speech samples to accomplish the recognition A test sample of any numeric in 1 to 5 is again recorded and then the algorithm is applied to recognise the same with recorded voices corresponding digits It has been found that a higher percentage (>90%) of recognition is reported when the threshold level is set to 025
TL;DR: Experiments carried out on a subset of the SRE 10 corpus using a scaled-down i-vector system indicate that direct DFT warping outperforms conventional MFCCs in most of the cases.
Abstract: Accuracy of speaker verification is high under controlled conditions but falls off rapidly in the presence of interfering sounds. This is because spectral features, such as Mel-frequency cepstral coefficients (MFCCs), are sensitive to additive noise. MFCCs are a particular realization of warped-frequency representation with low-frequency focus. But there are several alternative, potentially more robust, warped-frequency representations. We provide an experimental comparison of five warped-frequency features. They use exactly the same frequency warping function, the same number of coefficients and postprocessing, but differ in their internal computations. The compared variants are (1) conventional MFCCs from discrete Fourier transform (DFT), followed by Mel-scaled filterbank, (2) MFCCs via direct warping of DFT, followed by linear-scale filterbank, (3) warped linear prediction features, (4) perceptual minimum variance distortionless features and (5) recently proposed sparse Mel-scale histogram features. Experiments carried out on a subset of the SRE 10 corpus using a scaled-down i-vector system indicate that direct DFT warping outperforms conventional MFCCs in most of the cases. Index Terms: speaker recognition, noise, frequency warping
TL;DR: The application results indicate that two Wiener filter can achieve better results, in short stationary noise environments when face the low SNR.
Abstract: In the field of military communications, the SNR of speech signal is low because of the additive noise,. The two-level Wiener filter is adopted to analysis the noise characteristics. The DSP system architecture solutions are given out in this paper.The algorithms of the PSD mean module, Mel scale filter bank module, voice activity detection module and gain control module are introduced in detail. Finally, the application results are given out, which indicate that two Wiener filter can achieve better results, in short stationary noise environments when face the low SNR.
TL;DR: This paper shows how the macro and micro mechanical model is used in ASR tasks, and proposes a new approach that considers a new form to construct the bank filter in the authors' parametric representation.
Abstract: Recently the parametric representation using cochlea behavior has been used in different studies related with Automatic Speech Recognition (ASR). That is because this important organ of the hearing in the mammalians is the principal element used to make a transduction of the sound pressure that is received by the ear. In this paper we show how the macro and micro mechanical model is used in ASR tasks. We used the values that Neely founded in his work, related with the macro and micro mechanical model, such as was named, to set the central frequencies of a bank filter to obtain parameters from the speech used in a similar form as MFCC were constructed. We propose a new approach that considers a new form to construct the bank filter in our parametric representation. Then we used this distribution of the bank filter to have a new representation of the speech in frequency domain. It is important indicate that MFCC parameters use Mel scale to create a bank filter where central frequencies of each filter is in function of the scale mentioned above. We used the response of the Neely's model behavior to create the central frequencies of the bank filter mentioned above, then we substitute the Mel scale function by another representation. We use the place theory, and we reach a 98.5% of performance, for a task that uses isolated digits pronounced by 5 different speakers. Neely's model was used because a set of parameters of the cochlea as mass, damping and stiffness, among others, when are substituted inside the model make the response obtained is closer than von Bekesy proposed in his preliminary work about principle function of the cochlea.