Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
- Authors
- Lim, DongHoon; Kim, YoungChae; Kim, Dong-Hyun; Yang, Da-Hee; Chang, Joon-Hyuk
- Issue Date
- Apr-2026
- Publisher
- Institute of Electrical and Electronics Engineers Inc.
- Keywords
- Audio-Visual Speech Recognition; Cross-Modal Fusion; Noise-Robust ASR; Router-Gated Cross Attention
- Citation
- ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop, pp 1 - 7
- Pages
- 7
- Indexed
- SCOPUS
- Journal Title
- ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
- Start Page
- 1
- End Page
- 7
- URI
- https://scholarworks.bwise.kr/hanyang/handle/2021.sw.hanyang/219640
- DOI
- 10.1109/ASRU65441.2025.11434748
- ISSN
- 2997-6928
2997-6995
- Abstract
- Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise.
- Files in This Item
-
Go to Link
- Appears in
Collections - 서울 공과대학 > 서울 융합전자공학부 > 1. Journal Articles

Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.