Skip to main navigation Skip to search Skip to main content

Multi-modal query expansion for video object instances retrieval

  • Institut Mines-Télécom

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper we tackle the issue of object instances retrieval in video repositories using minimum information from the user (e.g., textual description/tags). Starting for a set of tags, images containing the object of interest are crawled from popular image search engines and repositories (e.g., Bing1, Fickr2, Google3) and the positive and most representative instances of the object are automatically identified. These positive images are then used to generate a visual query descriptor and to retrieve videos containing the object of the interest. This multi-modal approach makes it possible to retrieve video content through images obtained from textual queries, without the use of any advanced learning technique. We test out method on the Flickr corpus of the TRECVID 2012 Instance Search Task.

Original languageEnglish
Title of host publicationProceedings of the 13th IAPR International Conference on Machine Vision Applications, MVA 2013
PublisherMVA Organization
Pages214-217
Number of pages4
ISBN (Print)9784901122139
Publication statusPublished - 1 Jan 2013
Externally publishedYes
Event13th IAPR International Conference on Machine Vision Applications, MVA 2013 - Kyoto, Japan
Duration: 20 May 201323 May 2013

Publication series

NameProceedings of the 13th IAPR International Conference on Machine Vision Applications, MVA 2013

Conference

Conference13th IAPR International Conference on Machine Vision Applications, MVA 2013
Country/TerritoryJapan
CityKyoto
Period20/05/1323/05/13

Fingerprint

Dive into the research topics of 'Multi-modal query expansion for video object instances retrieval'. Together they form a unique fingerprint.

Cite this