Passer à la navigation principale Passer à la recherche Passer au contenu principal

Active Speaker Recognition using Cross Attention Audio-Video Fusion

  • ARTEMIS Department
  • Institut Polytechnique de Paris
  • University 'Politehnica' of Bucharest

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

5 Citations (Scopus)

Résumé

The audio-video based multimodal active speaker recognition from video streams has attracted the attention of the scientific community due to its wide range of applications, such as human centered computing or semantic video understanding. Most of the existing techniques use early or late fusion audio- video (A-V) strategies without considering completely the inter- modal and intra-modal interactions. In this context, this research work proposes a novel cross-modal attention mechanism based on visual and audio modalities designed to capture the complex spatiotemporal relationship between descriptors and to fuse complementary information from multiple modalities. First, we perform the representation learning of audio and video using deep convolutional neural networks (CNNs). Secondly, we feed the features of both modalities to a cross attention block by fusing A-V features at the model level. Finally, we obtain the identity of the active speaker and associate to each character the corresponding subtitle segment. The experimental evaluation performed on 30 videos validates the approach with average F1-scores superior to 88%. The effectiveness of the proposed system architecture is compared against state-of-the-art methods and demonstrates accuracy gains of more than 3%.

langue originaleAnglais
titre2022 10th European Workshop on Visual Information Processing, EUVIP 2022 - Proceedings
EditeurInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronique)9781665466233
Les DOIs
étatPublié - 1 janv. 2022
Evénement10th European Workshop on Visual Information Processing, EUVIP 2022 - Lisbon, Portugal
Durée: 11 sept. 202214 sept. 2022

Série de publications

NomProceedings - European Workshop on Visual Information Processing, EUVIP
Volume2022-September
ISSN (imprimé)2471-8963

Une conférence

Une conférence10th European Workshop on Visual Information Processing, EUVIP 2022
Pays/TerritoirePortugal
La villeLisbon
période11/09/2214/09/22

Empreinte digitale

Examiner les sujets de recherche de « Active Speaker Recognition using Cross Attention Audio-Video Fusion ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation