𝔖 Bobbio Scriptorium
✦   LIBER   ✦

Investigation on LP-residual representations for speaker identification

✍ Scribed by M. Chetouani; M. Faundez-Zanuy; B. Gas; J.L. Zarader


Book ID
104077311
Publisher
Elsevier Science
Year
2009
Tongue
English
Weight
231 KB
Volume
42
Category
Article
ISSN
0031-3203

No coin nor oath required. For personal study only.

✦ Synopsis


Feature extraction is an essential and important step for speaker recognition systems. In this paper, we propose to improve these systems by exploiting both conventional features such as mel frequency cepstral coding (MFCC), linear predictive cepstral coding (LPCC) and non-conventional ones. The method exploits information present in the linear predictive (LP) residual signal. The features extracted from the LP-residue are then combined to the MFCC or the LPCC. We investigate two approaches termed as temporal and frequential representations. The first one consists of an auto-regressive (AR) modelling of the signal followed by a cepstral transformation in a similar way to the LPC-LPCC transformation. In order to take into account the non-linear nature of the speech signals we used two estimation methods based on second and third-order statistics. They are, respectively, termed as R-SOS-LPCC (residual plus secondorder statistic based estimation of the AR model plus cepstral transformation) and R-HOS-LPCC (higher order). Concerning the frequential approach, we exploit a filter bank method called the power difference of spectra in sub-band (PDSS) which measures the spectral flatness over the sub-bands. The resulting features are named R-PDSS. The analysis of these proposed schemes are done over a speaker identification problem with two different databases. The first one is the Gaudi database and contains 49 speakers. The main interest lies in the controlled acquisition conditions: mismatch between the microphones and the interval sessions. The second database is the well-known NTIMIT corpus with 630 speakers. The performances of the features are confirmed over this larger corpus. In addition, we propose to compare traditional features and residual ones by the fusion of recognizers (feature extractor + classifier). The results show that residual features carry speaker-dependent features and the combination with the LPCC or the MFCC shows global improvements in terms of robustness under different mismatches. A comparison between the residual features under the opinion fusion framework gives us useful information about the potential of both temporal and frequential representations.


πŸ“œ SIMILAR VOLUMES


Identification of single amino acid resi
✍ Dr O. Reyes; M. G. Vallespi; H. E. Garay; L. J. Cruz; L. J. GonzΓ‘lez; G. Chinea; πŸ“‚ Article πŸ“… 2002 πŸ› John Wiley and Sons 🌐 English βš– 112 KB

## Abstract Lipopolysaccharide binding protein (LBP) is a 60 kDa acute phase glycoprotein capable of binding to LPS of Gram‐negative bacteria and facilitating its interaction with cellular receptors. This process is thought to be of great importance in systemic inflammatory reactions such as septic