In real acoustic environments, audio signals are typically composed of multiple speakers, various environmental sound events, background noise, and reverberation. Traditional audio processing methods mostly focus on recognition, classification, separation, or enhancement of entire audio segments, making it difficult to meet users' needs for precise acquisition of specific target sounds. With the development of applications such as intelligent conferencing, human-computer interaction, robot audition, hearing enhancement, security monitoring, and multimedia content understanding, how to extract designated target sounds on demand from complex mixed audio has become an important research problem in the field of intelligent audio processing.
This session focuses on the core direction of "target sound extraction," with emphasis on how to utilize reference audio, target speech characteristics, text descriptions, visual cues, or other prior conditions to extract user-specified target speech, target sound events, or specific sound sources from complex mixed audio. Unlike traditional speech separation, sound event detection, or audio enhancement tasks, target sound extraction emphasizes using target conditions as constraints to selectively extract specific sounds of interest from complex mixed audio, achieving selective extraction and on-demand acquisition of target sounds.
The purpose of establishing this session is to promote the transition of audio processing from general analysis of overall signals to conditional and fine-grained extraction oriented toward target objects. Target sound extraction not only focuses on "whether sounds can be separated," but more importantly on "whether designated sounds can be extracted according to user intent." This requires models to have effective understanding capabilities for target conditions, modeling capabilities for complex mixed signals, and robust generalization capabilities in noisy, reverberant, multi-speaker, multi-source overlapping, and open-category environments. Therefore, this session helps promote cross-disciplinary integration of speech processing, sound source separation, sound event analysis, multimodal learning, and audio large models, and advances the establishment of a more unified target sound extraction methodology system.
This session has important academic value and application significance. From an academic perspective, target sound extraction involves frontier problems such as target condition representation, mixed audio modeling, cross-modal target prompting, unified extraction of speech and non-speech sounds, open-category generalization, and low-latency robust processing, helping to expand the research boundaries of traditional audio separation and enhancement tasks. From an application perspective, this technology can be widely applied to scenarios such as designated speaker speech extraction in intelligent conferences, target speech enhancement in hearing aids, user speech acquisition in human-computer interaction, key sound source extraction in robot audition, abnormal sound acquisition in security monitoring, and target sound separation in multimedia content editing, which is of great significance for improving the perception capabilities, interaction capabilities, and practical value of intelligent audio systems in real complex environments.
Shuai Wang (Nanjing University): shuaiwang@nju.edu.cn
Ming Li (The Chinese University of Hong Kong, Shenzhen): mingli369@cuhk.edu.cn
Xinyuan Qian (University of Science and Technology Beijing): qianxy@ustb.edu.cn
Jian Guan (Harbin Engineering University): j.guan@hrbeu.edu.cn