A Momentum-Based Framework with Contrastive Data Generation for Robust Sound Source Localization
- Authors
- Kim, Hyun-Soo; Yang, Da-Hee; Chang, Joon-Hyuk
- Issue Date
- Apr-2026
- Publisher
- Institute of Electrical and Electronics Engineers Inc.
- Keywords
- contrastive learning; curriculum learning; exponentially moving average; momentum; sound source localization
- Citation
- ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop, pp 1 - 7
- Pages
- 7
- Indexed
- SCOPUS
- Journal Title
- ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
- Start Page
- 1
- End Page
- 7
- URI
- https://scholarworks.bwise.kr/hanyang/handle/2021.sw.hanyang/219344
- DOI
- 10.1109/ASRU65441.2025.11434773
- ISSN
- 2997-6928
2997-6995
- Abstract
- We propose MoCo-SSL, a momentum-based contrastive learning framework for multi-channel sound source localization (SSL) that enhances azimuth-aware representation learning. While prior SSL studies have used contrastive learning to handle varied acoustic conditions, we emphasize hard negatives-pairs with distinct azimuths recorded in the same room-for learning fine-grained spatial cues. A curriculum-based strategy gradually increases the proportion of such samples to raise task difficulty. The momentum contrast design employs a key encoder that maintains stable embeddings during curriculum transitions and receives audio with less noise and reverberation to produce clearer azimuth cues, thereby guiding the query encoder toward robust representations. Experiments show that MoCo-SSL consistently surpasses baselines, demonstrating the value of structured and noise-resilient representation learning in challenging SSL scenarios.
- Files in This Item
-
Go to Link
- Appears in
Collections - 서울 공과대학 > 서울 융합전자공학부 > 1. Journal Articles

Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.