Passer à la navigation principale Passer à la recherche Passer au contenu principal

Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

  • Thomas Serre
  • , Mathieu Fontaine
  • , Éric Benhaim
  • , Slim Essid
  • Orosound
  • Institut Polytechnique de Paris

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding's quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-the-fly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load.

langue originaleAnglais
titre2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Proceedings
rédacteurs en chefBhaskar D Rao, Isabel Trancoso, Gaurav Sharma, Neelesh B. Mehta
EditeurInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronique)9798350368741
Les DOIs
étatPublié - 1 janv. 2025
Evénement2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Hyderabad, Inde
Durée: 6 avr. 202511 avr. 2025

Série de publications

NomICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
ISSN (imprimé)1520-6149

Une conférence

Une conférence2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025
Pays/TerritoireInde
La villeHyderabad
période6/04/2511/04/25

Empreinte digitale

Examiner les sujets de recherche de « Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation