Passer à la navigation principale Passer à la recherche Passer au contenu principal

Conditional independence for pretext task selection in self-supervised speech representation learning

  • Salah Zaiem
  • , Titouan Parcollet
  • , Slim Essid
  • Institut Polytechnique de Paris
  • Avignon Université

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Through solving pretext tasks, self-supervised learning (SSL) leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstream task. A common pretext task consists in pretraining a SSL model on pseudo-labels derived from the original signal. This technique is particularly relevant for speech data where various meaningful signal processing features may serve as pseudolabels. However, the process of selecting pseudo-labels, for speech or other types of data, remains mostly unexplored and currently relies on observing the results on the final downstream task. Nevertheless, this methodology is not sustainable at scale due to substantial computational (hence carbon) costs. Thus, this paper introduces a practical and theoretical framework to select relevant pseudo-labels with respect to a given downstream task. More precisely, we propose a functional estimator of the pseudo-label utility grounded in the conditional independence theory, which does not require any training. The experiments conducted on speaker recognition and automatic speech recognition validate our estimator, showing a significant correlation between the performance observed on the downstream task and the utility estimates obtained with our approach, facilitating the prospection of relevant pseudo-labels for selfsupervised speech representation learning.

langue originaleAnglais
titre22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021
EditeurInternational Speech Communication Association
Pages1281-1285
Nombre de pages5
ISBN (Electronique)9781713836902
Les DOIs
étatPublié - 1 janv. 2021
Evénement22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021 - Brno, République tchcque
Durée: 30 août 20213 sept. 2021

Série de publications

NomProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Volume2
ISSN (imprimé)2308-457X
ISSN (Electronique)2958-1796

Une conférence

Une conférence22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021
Pays/TerritoireRépublique tchcque
La villeBrno
période30/08/213/09/21

Empreinte digitale

Examiner les sujets de recherche de « Conditional independence for pretext task selection in self-supervised speech representation learning ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation