Skip to main navigation Skip to search Skip to main content

Lip Reading Across Languages: A Cross-Modal Framework Leveraging Foundation Models

  • University 'Politehnica' of Bucharest
  • Telecom Sudparis

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Visual Speech Recognition (VSR), or lip reading, is essential in scenarios where audio signals are absent or degraded. In this paper, we introduce an end-to-end framework that integrates visual representations into a pretrained Large Language Model (LLM), enabling transcription that leverages multimodal context for improved accuracy and robustness. The core of our method is a cross-modal attention module that establishes fine-grained alignment between audio and visual streams during training, paired with lightweight adapters for seamless multimodal integration. At inference, the model relies solely on visual data, benefiting from audio-guided learning to enhance transcription accuracy. The proposed framework enables robust adaptation across varied linguistic conditions, yielding superior generalization and performance. Our experiments across Latin-script languages demonstrate consistent improvements over the current state of the art, yielding 1.53%-3.83% absolute reductions in WER. Experiments on Romanian, a previously unseen language, reveal strong zero-shot generalization and significant improvements after fine-tunning.

Original languageEnglish
Title of host publicationCBMI 2025 - 2025 International Conference on Content-Based Multimedia Indexing, Conference Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331555009
DOIs
Publication statusPublished - 1 Jan 2025
Event2025 International Conference on Content-Based Multimedia Indexing, CBMI 2025 - Dublin, Ireland
Duration: 22 Oct 202524 Oct 2025

Publication series

NameCBMI 2025 - 2025 International Conference on Content-Based Multimedia Indexing, Conference Proceedings

Conference

Conference2025 International Conference on Content-Based Multimedia Indexing, CBMI 2025
Country/TerritoryIreland
CityDublin
Period22/10/2524/10/25

Keywords

  • cross modal attention
  • large language model
  • lip reading

Fingerprint

Dive into the research topics of 'Lip Reading Across Languages: A Cross-Modal Framework Leveraging Foundation Models'. Together they form a unique fingerprint.

Cite this