Xurong Xie is currently an Associate Researcher at the Human-Computer Interaction Technology and Intelligent Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences. He received his Ph.D. from the Department of Electronic Engineering, The Chinese University of Hong Kong, and has worked and studied at the Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, and the Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong. His research interests include disordered speech processing, adaptive technology, speech and language modeling, and neural speech decoding. He has published over 60 journal and conference papers in related fields. His proposed Bayesian adaptive technique for neural network acoustic models won the ICASSP 2019 Best Student Paper Award. His paper on speech foundation model-based disordered speech recognition and detection was selected as one of the IEEE SPS 2024–2025 TASLP Journal Top 25 Downloaded Papers. He also received awards in the ISCSLP 2024 Multimodal Dysarthria Severity Assessment Challenge and the HHME Outstanding Paper Award.
He has undertaken national and CAS research tasks including the NSFC Youth Program, sub-project of the New Generation Artificial Intelligence Major Project, sub-project of the National Key R&D Program, sub-project of the ISCAS Major Project, and CAS Youth Innovation Promotion Association Membership Project.
Automatic speech recognition (ASR) technology for normal speech has made rapid progress in recent decades, but accurate recognition of dysarthric and cognitively impaired speech remains a highly challenging task. Dysarthria is a common speech motor disorder often caused by motor control abnormalities due to cerebral palsy, amyotrophic lateral sclerosis, and stroke; neurocognitive disorders such as Alzheimer's disease are also very common in elderly populations with speech-language function impairment.
ASR technology customized for users with impairments can not only significantly improve their quality of life and assist rehabilitation, but also support automated early diagnosis of neurocognitive function impairment. Current ASR technology is mainly designed for normal speech of healthy non-elderly users, while dysarthric and cognitively impaired speech poses many challenges, including: significant acoustic mismatch with normal speech due to motor control abnormalities and aging processes; data scarcity caused by difficulties in data collection; and high heterogeneity at the speaker level.
Pre-trained speech foundation models have been widely applied to various ASR tasks. However, migrating them to impaired speech scenarios through data-intensive parameter fine-tuning faces dual challenges of domain data scarcity and distribution mismatch, easily leading to performance and generalization degradation. Therefore, dedicated system architectures, model representations, and training and modeling methods are needed to address these issues, such as model fusion methods for impaired speech, discrete token representations, semi-supervised learning methods, and adaptive modeling techniques.