Skip to main navigation Skip to search Skip to main content

A generic classification system for multi-channel audio indexing: Application to speech and music detection

  • Sorbonne Université

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Citations (Scopus)

Abstract

There is a rise in the number 3D audio-visual productions and archives that creates a need for indexation of 3D contents. Event detection using audio modality is a difficult task. The standard way to do classification on 3D audio is to first down-mix to mono audio and classify on that. In this paper, we describe a generic classifier for multi-channel audio event detection and propose several information fusion strategies. Our system is evaluated on a speech and music detection task on the audio of 3D movies. We improve the classification performances on our database by 1.5% for speech detection, and 8% for music detection, compared to the standard downmixing method. We also provide a comparison of several information fusion methods in the experiments.

Original languageEnglish
Title of host publication2013 14th International Workshop on Image Analysis for Multimedia Interactive Services, WIAMIS 2013
DOIs
Publication statusPublished - 13 Nov 2013
Event2013 14th International Workshop on Image Analysis for Multimedia Interactive Services, WIAMIS 2013 - Paris, France
Duration: 3 Jul 20135 Jul 2013

Publication series

NameInternational Workshop on Image Analysis for Multimedia Interactive Services
ISSN (Print)2158-5873
ISSN (Electronic)2158-5881

Conference

Conference2013 14th International Workshop on Image Analysis for Multimedia Interactive Services, WIAMIS 2013
Country/TerritoryFrance
CityParis
Period3/07/135/07/13

Fingerprint

Dive into the research topics of 'A generic classification system for multi-channel audio indexing: Application to speech and music detection'. Together they form a unique fingerprint.

Cite this