Accession Number:

ADA555288

Title:

Perturbation and Pitch Normalization as Enhancements to Speaker Recognition

Descriptive Note:

Conference paper

Corporate Author:

AIR FORCE RESEARCH LAB ROME NY

Report Date:

2009-04-01

Pagination or Media Count:

5.0

Abstract:

This study proposes an approach to improving speaker recognition through the process of minute vocal tract length perturbation of training files, coupled with pitch normalization for both train and test data. The notion of perturbation as a method for improving the robustness of training data for supervised classification is taken from the field of optical character recognition, where distorting characters within a certain range has shown strong improvements across disparate conditions. This paper demonstrates that acoustic perturbation, in this case analysis, distortion, and resynthesis of vocal tract length for a given speaker, significantly improves speaker recognition when the resulting files are used to augment or replace the training data. A pitch length normalization technique is also discussed, which is combined with perturbation to improve open-set speaker recognition from an EER of 20 to 6.7.

Subject Categories:

  • Acoustics
  • Voice Communications

Distribution Statement:

APPROVED FOR PUBLIC RELEASE