Passer à la navigation principale Passer à la recherche Passer au contenu principal

Robust visual features for the multimodal identification of unregistered speakers in TV talk-shows

  • CNRS LTCI
  • Research Department of Institut National de l'Audiovisuel

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

In this paper we propose a novel multimodal method for identifying unregistered speakers in a TV talk-show using a semi-supervised learning approach based on Support Vector Machines. Our study highlights the fact that specific visual features prove to be very efficient for this particular type of video content which is edited from multi-camera recordings. These visual features, motivated by prior knowledge on the approach followed by the TV director in choosing the appropriate shots, are found to bring a significant improvement in identification accuracy when used together with classic audio Mel-frequency cepstral coefficients (+8% compared to various baseline systems, in particular a standard audio only system).

langue originaleAnglais
titre2010 IEEE International Conference on Image Processing, ICIP 2010 - Proceedings
Pages1469-1472
Nombre de pages4
Les DOIs
étatPublié - 1 déc. 2010
Modification externeOui
Evénement2010 17th IEEE International Conference on Image Processing, ICIP 2010 - Hong Kong, Hong-Kong
Durée: 26 sept. 201029 sept. 2010

Série de publications

NomProceedings - International Conference on Image Processing, ICIP
ISSN (imprimé)1522-4880

Une conférence

Une conférence2010 17th IEEE International Conference on Image Processing, ICIP 2010
Pays/TerritoireHong-Kong
La villeHong Kong
période26/09/1029/09/10

Empreinte digitale

Examiner les sujets de recherche de « Robust visual features for the multimodal identification of unregistered speakers in TV talk-shows ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation