Cited 0 time in
Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Lim, DongHoon | - |
| dc.contributor.author | Kim, YoungChae | - |
| dc.contributor.author | Kim, Dong-Hyun | - |
| dc.contributor.author | Yang, Da-Hee | - |
| dc.contributor.author | Chang, Joon-Hyuk | - |
| dc.date.accessioned | 2026-07-24T07:00:16Z | - |
| dc.date.available | 2026-07-24T07:00:16Z | - |
| dc.date.issued | 2026-04 | - |
| dc.identifier.issn | 2997-6928 | - |
| dc.identifier.issn | 2997-6995 | - |
| dc.identifier.uri | https://scholarworks.bwise.kr/hanyang/handle/2021.sw.hanyang/219640 | - |
| dc.description.abstract | Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise. | - |
| dc.format.extent | 7 | - |
| dc.language | 영어 | - |
| dc.language.iso | ENG | - |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | - |
| dc.title | Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion | - |
| dc.type | Article | - |
| dc.identifier.doi | 10.1109/ASRU65441.2025.11434748 | - |
| dc.identifier.scopusid | 2-s2.0-105036549498 | - |
| dc.identifier.bibliographicCitation | ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop, pp 1 - 7 | - |
| dc.citation.title | ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop | - |
| dc.citation.startPage | 1 | - |
| dc.citation.endPage | 7 | - |
| dc.type.docType | Conference paper | - |
| dc.description.isOpenAccess | N | - |
| dc.description.journalRegisteredClass | scopus | - |
| dc.subject.keywordPlus | Audio acoustics | - |
| dc.subject.keywordPlus | Audio signal processing | - |
| dc.subject.keywordPlus | Audio systems | - |
| dc.subject.keywordPlus | Routers | - |
| dc.subject.keywordPlus | Sound reproduction | - |
| dc.subject.keywordPlus | Speech communication | - |
| dc.subject.keywordPlus | Speech recognition | - |
| dc.subject.keywordAuthor | Audio-Visual Speech Recognition | - |
| dc.subject.keywordAuthor | Cross-Modal Fusion | - |
| dc.subject.keywordAuthor | Noise-Robust ASR | - |
| dc.subject.keywordAuthor | Router-Gated Cross Attention | - |
| dc.identifier.url | https://ieeexplore.ieee.org/document/11434748 | - |
Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.
222, Wangsimni-ro, Seongdong-gu, Seoul, 04763, Korea+82-2-2220-1366
COPYRIGHT © 2024 HANYANG UNIVERSITY.
Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.
