With the rapid development of artificial intelligence, large models, and embodied intelligence technologies, speech interaction is gradually evolving from static, single-person, near-field environments to open, complex, and dynamic real-world scenarios. Emerging applications such as human-computer interaction, smart homes, intelligent cockpits, intelligent robots, and mixed reality impose higher demands on speech processing algorithms: not only the ability to "hear and hear clearly," but also accurate perception of dynamic acoustic environments, distinguishing and understanding spatial relationships among multiple sound sources, and achieving stable and robust speech perception capabilities in real-world scenarios.
As an important carrier for spatial acoustic information acquisition, microphone arrays can provide rich spatiotemporal information and serve as a fundamental basis for dynamic sound field modeling, spatial audio analysis, and robust speech perception. In recent years, traditional array signal processing has been continuously integrated with deep learning, large models, and generative artificial intelligence, achieving significant progress in high-precision sound field reconstruction, dynamic sound source perception, speech enhancement, and sound source separation. Meanwhile, with the development of self-supervised pre-training methods and multimodal foundation models, how to organically combine array observation information, spatial acoustic theory, and data-driven models to construct a new generation of spatial speech intelligence systems with environmental understanding capabilities has become an important research hotspot in the international speech and acoustics community.
This special session focuses on sound field modeling and speech perception based on microphone arrays in dynamic scenarios, addressing key issues such as dynamic spatial sound field modeling and simulation, multi-source acoustic perception, spatial audio generation, and robust speech processing. It emphasizes the integration of physical models and data-driven methods, multimodal information collaborative modeling, and the development of spatial audio foundation models, promoting deep cross-disciplinary integration among array signal processing, speech processing, artificial intelligence, and related application fields.
This session aims to build a platform for exchange and cooperation among experts and scholars in relevant domestic fields, promote the development of key technologies such as dynamic sound field modeling, spatial audio perception, and robust speech processing, and provide theoretical foundations and key technical support for a new generation of intelligent interaction applications including robots, intelligent terminals, intelligent meetings, digital humans, and embodied intelligence. The session holds important scientific significance and application value: at the scientific level, the "physical mechanism + data-driven" fusion paradigm explored in this session can effectively compensate for the shortcomings of deep learning methods in interpretability and physical consistency, opening new paths for the cross-disciplinary integration of computational acoustics and artificial intelligence; at the application level, these core technologies are of great significance for meeting urgent needs in key fields such as defense communications, smart homes, healthcare, and embodied intelligence, and can also provide valuable references for the development and continuous progress of related technologies in other fields.
Yanmin Qian (Shanghai Jiao Tong University): yanminqian@sjtu.edu.cn
Gongping Huang (Wuhan University): gongpinghuang@whu.edu.cn
Wangyou Zhang (Shanghai Jiao Tong University): wyz-97@sjtu.edu.cn
Zhongxin Bai (Wuhan University): 18392387962@163.com