Passer à la navigation principale Passer à la recherche Passer au contenu principal

Learning visual voice activity detection with an automatically annotated dataset

  • Sylvain Guy
  • , Stéphane Lathuilière
  • , Pablo Mesejo
  • , Radu Horaud
  • LTHE (UMR 5564 CNRS/IRD/Université de Grenoble)
  • University of Granada

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Visual voice activity detection (V-VAD) uses visual features to predict whether a person is speaking or not. V-VAD is useful whenever audio VAD (A-VAD) is inefficient either because the acoustic signal is difficult to analyze or because it is simply missing. We propose two deep architectures for V-VAD, one based on facial landmarks and one based on optical flow. Moreover, available datasets, used for learning and for testing V-VAD, lack content variability. We introduce a novel methodology to automatically create and annotate very large datasets in-the-wild - WildVVAD - based on combining A-VAD with face detection and tracking. A thorough empirical evaluation shows the advantage of training the proposed deep V-VAD models with this dataset.

langue originaleAnglais
titreProceedings of ICPR 2020 - 25th International Conference on Pattern Recognition
EditeurInstitute of Electrical and Electronics Engineers Inc.
Pages4851-4856
Nombre de pages6
ISBN (Electronique)9781728188089
Les DOIs
étatPublié - 1 janv. 2020
Evénement25th International Conference on Pattern Recognition, ICPR 2020 - Virtual, Online, Italie
Durée: 10 janv. 202115 janv. 2021

Série de publications

NomProceedings - International Conference on Pattern Recognition
ISSN (imprimé)1051-4651

Une conférence

Une conférence25th International Conference on Pattern Recognition, ICPR 2020
Pays/TerritoireItalie
La villeVirtual, Online
période10/01/2115/01/21

Empreinte digitale

Examiner les sujets de recherche de « Learning visual voice activity detection with an automatically annotated dataset ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation