TY - GEN
T1 - Lip Reading Across Languages
T2 - 2025 International Conference on Content-Based Multimedia Indexing, CBMI 2025
AU - Tapu, Ruxandra
AU - Mocanu, Bogdan
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/1/1
Y1 - 2025/1/1
N2 - Visual Speech Recognition (VSR), or lip reading, is essential in scenarios where audio signals are absent or degraded. In this paper, we introduce an end-to-end framework that integrates visual representations into a pretrained Large Language Model (LLM), enabling transcription that leverages multimodal context for improved accuracy and robustness. The core of our method is a cross-modal attention module that establishes fine-grained alignment between audio and visual streams during training, paired with lightweight adapters for seamless multimodal integration. At inference, the model relies solely on visual data, benefiting from audio-guided learning to enhance transcription accuracy. The proposed framework enables robust adaptation across varied linguistic conditions, yielding superior generalization and performance. Our experiments across Latin-script languages demonstrate consistent improvements over the current state of the art, yielding 1.53%-3.83% absolute reductions in WER. Experiments on Romanian, a previously unseen language, reveal strong zero-shot generalization and significant improvements after fine-tunning.
AB - Visual Speech Recognition (VSR), or lip reading, is essential in scenarios where audio signals are absent or degraded. In this paper, we introduce an end-to-end framework that integrates visual representations into a pretrained Large Language Model (LLM), enabling transcription that leverages multimodal context for improved accuracy and robustness. The core of our method is a cross-modal attention module that establishes fine-grained alignment between audio and visual streams during training, paired with lightweight adapters for seamless multimodal integration. At inference, the model relies solely on visual data, benefiting from audio-guided learning to enhance transcription accuracy. The proposed framework enables robust adaptation across varied linguistic conditions, yielding superior generalization and performance. Our experiments across Latin-script languages demonstrate consistent improvements over the current state of the art, yielding 1.53%-3.83% absolute reductions in WER. Experiments on Romanian, a previously unseen language, reveal strong zero-shot generalization and significant improvements after fine-tunning.
KW - cross modal attention
KW - large language model
KW - lip reading
UR - https://www.scopus.com/pages/publications/105033156348
U2 - 10.1109/CBMI66578.2025.11339323
DO - 10.1109/CBMI66578.2025.11339323
M3 - Conference contribution
AN - SCOPUS:105033156348
T3 - CBMI 2025 - 2025 International Conference on Content-Based Multimedia Indexing, Conference Proceedings
BT - CBMI 2025 - 2025 International Conference on Content-Based Multimedia Indexing, Conference Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 22 October 2025 through 24 October 2025
ER -