Skip to main navigation Skip to search Skip to main content

Video structuring: From pixels to visual entities

  • CNRS SAMOVAR UMR 5157

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Citation (Scopus)

Abstract

In this paper we propose a method for automatic structuring of video documents. The video is firstly segmented into shots based on a scale space filtering graph partition method. For each detected shot the associated static summary is developed using a leap key-frame extraction method. Based on the representative images obtained, we introduce next a combined spatial and temporal video attention model that is able to recognize moving salient objects. The proposed approach extends the state-of-the-art image region based contrast saliency with a temporal attention model. Different types of motion presented in the current shot are determined using a set of homographic transforms, estimated by recursively applying the RANSAC algorithm on the interest point correspondence. Finally, a decision is taken based on the combined spatial and temporal attention models. The experimental results validate the proposed framework and demonstrate that our approach is effective for various types of videos, including noisy and low resolution data.

Original languageEnglish
Title of host publicationProceedings of the 20th European Signal Processing Conference, EUSIPCO 2012
PublisherEuropean Signal Processing Conference, EUSIPCO
Pages1583-1587
Number of pages5
ISBN (Print)9781467310680
Publication statusPublished - 1 Jan 2012
Externally publishedYes
Event20th European Signal Processing Conference, EUSIPCO 2012 - Bucharest, Romania
Duration: 27 Aug 201231 Aug 2012

Publication series

NameEuropean Signal Processing Conference
ISSN (Print)2219-5491
ISSN (Electronic)2076-1465

Conference

Conference20th European Signal Processing Conference, EUSIPCO 2012
Country/TerritoryRomania
CityBucharest
Period27/08/1231/08/12

Keywords

  • RANSAC algorithm
  • Saliency maps
  • homography transform
  • temporal attention model

Fingerprint

Dive into the research topics of 'Video structuring: From pixels to visual entities'. Together they form a unique fingerprint.

Cite this