Passer à la navigation principale Passer à la recherche Passer au contenu principal

EARS4SEE: A multimodal audio description system dedicated to blind and visually impaired users

  • University 'Politehnica' of Bucharest

Résultats de recherche: Contribution à un journalArticleRevue par des pairs

Résumé

In recent years, automatic audio description (AD) generation has become an important research domain within accessibility and assistive technology, driven by its potential to enhance content understanding, social integration, and cognitive engagement for individuals with visual impairments (VI). In this paper, we introduce EARS4SEE, a novel multimodal framework for AD generation that integrates semantic video analysis, character tracking, and adaptive temporal segmentation to enhance contextual coherence and narrative fluency. The proposed system integrates multi-stream fusion strategy, leveraging visual, textual, and audio modalities for character-centric, semantically enriched AD. Textual descriptions are synthesized into natural-sounding speech using state-of-the-art text-to-speech (TTS) techniques for an immersive experience. A core contribution of the proposed methodology involves the tracking-based character recognition module, which ensures temporally consistent character identification using an adaptive temporal attention mechanism. The approach mitigates inconsistencies from motion blur, occlusions, and scale variations, improving referential continuity. Additionally, EARS4SEE introduces an automated multimodal video segmentation pipeline, capturing long-range temporal dependencies to improve scene boundary detection and contextual alignment. The experimental evaluation carried out on the MAD-Eval-Named and TV-AD datasets validates the effectiveness of the proposed methodology, which leads to average CIDEr and LLM-AD-eval scores of 24.1 and 3.02, respectively. In addition, when compared to state-of-the-art techniques, the proposed architecture shows superior performances in terms of the CIDEr, with gains in accuracy ranging in the [1.72%, 10.2%] interval and an 8% increase in LLM-AD-eval scores.

langue originaleAnglais
Numéro d'article104772
journalComputer Vision and Image Understanding
Volume268
Les DOIs
étatPublié - 1 mai 2026

Empreinte digitale

Examiner les sujets de recherche de « EARS4SEE: A multimodal audio description system dedicated to blind and visually impaired users ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation