Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach
Researchers present a novel framework for speaker verification in non-verbal vocalizations (NVVs) like laughter and sighs, combining Data2Vec features with ECAPA-TDNN and a Mixture of Experts module. The approach reduces speech-to-NVV error rates from 38.93% to 22.66% while maintaining speech verification accuracy, addressing a critical gap in voice authentication systems as TTS and voice conversion technologies become increasingly sophisticated.