Skip to main navigation Skip to search Skip to main content

Automatic Data Augmentation for Domain Adapted Fine-Tuning of Self-Supervised Speech Representations

  • Salah Zaiem
  • , Titouan Parcollet
  • , Slim Essid
  • Institut Polytechnique de Paris
  • Samsung AI Center - Cambridge
  • University of Cambridge

Research output: Contribution to journalConference articlepeer-review

2 Citations (Scopus)

Abstract

Self-Supervised Learning (SSL) has allowed leveraging large amounts of unlabeled speech data to improve the performance of speech recognition models even with small annotated datasets. Despite this, speech SSL representations may fail while facing an acoustic mismatch between the pretraining and target datasets. To address this issue, we propose a novel supervised domain adaptation method, designed for cases exhibiting such a mismatch in acoustic domains. It consists in applying properly calibrated data augmentations on a large clean dataset, bringing it closer to the target domain, and using it as part of an initial fine-tuning stage. Augmentations are automatically selected through the minimization of a conditional-dependence estimator, based on the target dataset. The approach is validated during an oracle experiment with controlled distortions and on two amateur-collected low-resource domains, reaching better performances compared to the baselines in both cases.

Original languageEnglish
Pages (from-to)67-71
Number of pages5
JournalProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Volume2023-August
DOIs
Publication statusPublished - 1 Jan 2023
Event24th Annual conference of the International Speech Communication Association, Interspeech 2023 - Dublin, Ireland
Duration: 20 Aug 202324 Aug 2023

Keywords

  • domain adaptation
  • self-supervised learning

Fingerprint

Dive into the research topics of 'Automatic Data Augmentation for Domain Adapted Fine-Tuning of Self-Supervised Speech Representations'. Together they form a unique fingerprint.

Cite this