Tutorials

Advances in Speech Enhancement: From Predictive Estimation to Generative Modeling
Release Time:2026/9/30 22:20:19
Wenbin Jiang
Photo
Wenbin Jiang (江文斌)
Hangzhou Dianzi University · Distinguished Associate Professor, Master's Supervisor

Wenbin Jiang is a Distinguished Associate Professor and Master's Supervisor at the School of Communication Engineering, Hangzhou Dianzi University. He received his Ph.D. from the Department of Electronic Engineering at Shanghai Jiao Tong University and completed postdoctoral research in the Department of Computer Science and Engineering. His research focuses on intelligent speech information processing, including speech enhancement, speech wake-up, speech recognition, large speech understanding models, low-bitrate speech coding and decoding, and brain-inspired hearing. He has published more than twenty papers in top journals and conferences in the acoustics and speech field, such as IEEE/ACM TASLP, IEEE SPL, ICASSP, and Interspeech, and has been granted more than ten national invention patents. He has led projects including the Ministry of Science and Technology Yangtze River Delta Science and Technology Innovation Joint Research Program, the Zhejiang Provincial Natural Science Foundation, and multiple industry-university collaborative projects. He has long served as a reviewer for journals and conferences such as IEEE TASLP, ICASSP, Interspeech, IEEE SLT, and NCMMSC.

Speech enhancement research aims to remove interfering noise from noisy speech and is an important technical means of improving speech quality and intelligibility. In recent years, deep learning-based predictive methods (such as direct regression and mask estimation) have dominated the field. Although they have achieved significant breakthroughs in objective metrics, they tend to introduce perceptual distortion and perform limitedly in data reconstruction tasks such as bandwidth extension and codec-distortion restoration. Meanwhile, generative speech enhancement techniques represented by diffusion models and flow matching have developed rapidly. By explicitly modeling the prior of clean speech, they demonstrate superior perceptual quality and stronger robustness, and have attracted increasing attention from academia and industry. Therefore, how to understand the complementary relationship between the predictive and generative paradigms from a unified perspective has become a key issue in current speech enhancement research. This report systematically reviews single-channel speech enhancement and provides a unified analysis of the predictive and generative paradigms.

The content covers problem formulation, speech representation, the evolution of modeling paradigms (predictive, generative, and their hybrid architectures), and the evolution of model architectures, while also exploring emerging directions such as universal speech enhancement driven by speech foundation models and large language models. In addition, this report presents a comparative analysis of typical predictive, generative, and predictive-generative hybrid speech enhancement methods on representative benchmark datasets. Through this report, we hope to help the audience build a holistic understanding of the two paradigms and provide a reference for related research and applications.