Passer à la navigation principale Passer à la recherche Passer au contenu principal

Lip Reading Across Languages: A Cross-Modal Framework Leveraging Foundation Models

  • University 'Politehnica' of Bucharest
  • Telecom Sudparis

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Visual Speech Recognition (VSR), or lip reading, is essential in scenarios where audio signals are absent or degraded. In this paper, we introduce an end-to-end framework that integrates visual representations into a pretrained Large Language Model (LLM), enabling transcription that leverages multimodal context for improved accuracy and robustness. The core of our method is a cross-modal attention module that establishes fine-grained alignment between audio and visual streams during training, paired with lightweight adapters for seamless multimodal integration. At inference, the model relies solely on visual data, benefiting from audio-guided learning to enhance transcription accuracy. The proposed framework enables robust adaptation across varied linguistic conditions, yielding superior generalization and performance. Our experiments across Latin-script languages demonstrate consistent improvements over the current state of the art, yielding 1.53%-3.83% absolute reductions in WER. Experiments on Romanian, a previously unseen language, reveal strong zero-shot generalization and significant improvements after fine-tunning.

langue originaleAnglais
titreCBMI 2025 - 2025 International Conference on Content-Based Multimedia Indexing, Conference Proceedings
EditeurInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronique)9798331555009
Les DOIs
étatPublié - 1 janv. 2025
Evénement2025 International Conference on Content-Based Multimedia Indexing, CBMI 2025 - Dublin, Irlande
Durée: 22 oct. 202524 oct. 2025

Série de publications

NomCBMI 2025 - 2025 International Conference on Content-Based Multimedia Indexing, Conference Proceedings

Une conférence

Une conférence2025 International Conference on Content-Based Multimedia Indexing, CBMI 2025
Pays/TerritoireIrlande
La villeDublin
période22/10/2524/10/25

Empreinte digitale

Examiner les sujets de recherche de « Lip Reading Across Languages: A Cross-Modal Framework Leveraging Foundation Models ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation