Abstract
In recent years, automatic audio description (AD) generation has become an important research domain within accessibility and assistive technology, driven by its potential to enhance content understanding, social integration, and cognitive engagement for individuals with visual impairments (VI). In this paper, we introduce EARS4SEE, a novel multimodal framework for AD generation that integrates semantic video analysis, character tracking, and adaptive temporal segmentation to enhance contextual coherence and narrative fluency. The proposed system integrates multi-stream fusion strategy, leveraging visual, textual, and audio modalities for character-centric, semantically enriched AD. Textual descriptions are synthesized into natural-sounding speech using state-of-the-art text-to-speech (TTS) techniques for an immersive experience. A core contribution of the proposed methodology involves the tracking-based character recognition module, which ensures temporally consistent character identification using an adaptive temporal attention mechanism. The approach mitigates inconsistencies from motion blur, occlusions, and scale variations, improving referential continuity. Additionally, EARS4SEE introduces an automated multimodal video segmentation pipeline, capturing long-range temporal dependencies to improve scene boundary detection and contextual alignment. The experimental evaluation carried out on the MAD-Eval-Named and TV-AD datasets validates the effectiveness of the proposed methodology, which leads to average CIDEr and LLM-AD-eval scores of 24.1 and 3.02, respectively. In addition, when compared to state-of-the-art techniques, the proposed architecture shows superior performances in terms of the CIDEr, with gains in accuracy ranging in the [1.72%, 10.2%] interval and an 8% increase in LLM-AD-eval scores.
| Original language | English |
|---|---|
| Article number | 104772 |
| Journal | Computer Vision and Image Understanding |
| Volume | 268 |
| DOIs | |
| Publication status | Published - 1 May 2026 |
Keywords
- Adaptive narration
- Character recognition
- Large language models
- Multimodal audio description
- Scene segmentation
- Temporal attention
- Vision-language models
Fingerprint
Dive into the research topics of 'EARS4SEE: A multimodal audio description system dedicated to blind and visually impaired users'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver