Xinyuan Qian is an Associate Professor at the School of Computer and Communication Engineering, University of Science and Technology Beijing. Her main research interests include intelligent speech technology and multimodal (audio-visual) human-computer interaction. She received her Ph.D. in Computer Science from Queen Mary University of London, UK, and was a Postdoctoral Research Fellow at the National University of Singapore from 2020 to 2022. She has also conducted research at FBK in Italy and the Chinese University of Hong Kong (Shenzhen). She has published over 80 papers in top international conferences and journals, and serves as a Program Committee Member/Guest Editor for multiple international conferences/journals (such as IROS, ICASSP, ECAI, TCE, etc.). She has led or participated in projects including the National Natural Science Foundation of China Youth Project, Beijing Natural Science Foundation General Project, Science and Technology Innovation 2030 — New Generation Artificial Intelligence Major Project, Beijing Natural Science Foundation Joint Project, CCF-Tencent Rhino-Bird Project, Shenzhen Institute of Big Data Research Project, Singapore Human-Computer Interaction Project, and Huawei Technology Cooperation Project. She was awarded the ACM Rising Star Award (Beijing), Beijing "High-Innovation Plan" Young Talent, and Beijing Image and Graphics Society Most Beautiful Female Scientist.
Real-world human-computer interaction scenarios are typically filled with reverberation, strong noise, and multi-speaker interference. Relying solely on single-modal audio signals for spatial auditory perception faces significant robustness challenges. This tutorial focuses on the cutting-edge direction of "Spatial Multimodal Auditory Perception," systematically exploring how to organically integrate visual spatial information (such as scene geometry and facial dynamics captured by cameras) with multi-microphone array audio signals to break through the performance bottlenecks of traditional machine hearing.
First, we will elaborate on the latest advances in 3D Sound Source Localization (3D SSL) in complex acoustic environments, with a focus on how audio-visual deep feature cross-attention mechanisms address localization robustness under non-line-of-sight and strong occlusion conditions. Second, this tutorial will deeply analyze multimodal Target Speaker Extraction (TSE) technology, explaining how to leverage visual cues to precisely lock onto and separate specific speakers in audio. Finally, we will discuss the future evolution trends of 3D spatial audio-visual reasoning in conjunction with national strategic planning and the needs of multimodal human-computer interaction. Through this tutorial, attendees will build a complete knowledge architecture spanning from physical modeling, cross-modal representation learning, to real-world deployment.