Young Researchers Forum

Speech Enhancement Methods Inspired by Auditory Perception Mechanisms
Release Time:2026/8/21 11:58:23
NCMMSC 2026 Young Researchers Forum - Nan Li
Nan Li
Photo
Nan Li
School of Computer and Information Engineering, Tianjin Normal University · Lecturer

Nan Li is a Lecturer at the School of Computer and Information Engineering, Tianjin Normal University, and holds a Ph.D. in Computer Science and Technology from Tianjin University. She pursued a dual degree at Japan Advanced Institute of Science and Technology (JAIST) and was funded by the China Scholarship Council for joint training at the National University of Singapore during her doctoral studies. Her main research interests include speech enhancement, speech endpoint detection, far-field speech recognition, speech dereverberation, and auditory perception computation, with a focus on auditory mechanism-inspired low-complexity speech processing methods.

She has published 16 papers in journals and conferences including Expert Systems with Applications, Speech Communication, ICASSP, and INTERSPEECH, with 7 first-author papers and 1 corresponding-author paper, and has applied for 5 invention patents as the primary inventor. She has received honors including Top 2 in the Intel Neuromorphic Deep Noise Suppression Challenge, Top 4 in the Audio Deep Synthesis Detection Challenge, and the Algorithm Elite Award in the Dialect Recognition Challenge.

Background noise in complex acoustic environments degrades speech quality and intelligibility, and affects the user experience of speech recognition, human-computer interaction, and hearing aid devices. This talk introduces a series of speech enhancement methods inspired by auditory perception mechanisms, centered on how the human auditory system perceives and processes noisy speech.

First, the talk explores how to simulate auditory encoding, auditory masking, and context-aware mechanisms to achieve speech endpoint detection in complex noise environments. Then, it introduces a dual-stream information perception framework for noise and speech, discussing how to use prior information about environmental noise and target speech to assist speech enhancement. Finally, it presents a harmonic-compensation auditory perception network combining cochlear frequency selectivity and pitch perception mechanisms, balancing speech quality, auditory naturalness, and model efficiency.

This talk will focus on how to transform physiological auditory mechanisms into learnable computational models, and the application prospects of bio-inspired methods in smart terminals, voice communication, and hearing aid devices.