Tutorials

Robust Dialogue Intelligence: Perception and Expression of Human-Computer Voice Interaction in Complex Scenarios
Release Time:2026/8/5 14:04:23
Rui Liu
Photo
Rui Liu (刘瑞)
Inner Mongolia University · Professor, Ph.D. Supervisor

Rui Liu is the Vice Dean of the School of Artificial Intelligence at Inner Mongolia University, Professor, and Ph.D. Supervisor. He was selected for the 10th China Association for Science and Technology Young Talent Lifting (Qingtuo) Program. He has led more than 10 national/provincial-level projects, including the National Natural Science Foundation of China General Program and Youth Program, Inner Mongolia Autonomous Region Outstanding Young Scientist Fund, Central Government Guiding Local Science and Technology Development Plan Project, Inner Mongolia Autonomous Region Key R&D and Achievement Transformation Plan Project, and Inner Mongolia Autonomous Region Grassland Talent Program. His main research direction is multilingual multimodal empathetic human-computer voice interaction. His related achievements have been published as first or corresponding author in top-tier academic conferences or journals such as IEEE-TASLP, IEEE-TAFFC, ACL, ACM MM, and AAAI, including 1 ESI highly cited paper and 2 IEEE international conference best paper awards. He received the First Prize of Inner Mongolia Autonomous Region Science and Technology Progress Award and the Second Prize of China Invention Association Invention and Entrepreneurship Award. He serves as Associate Editor for international journals Information Fusion, IEEE-TAFFC, and ACM-TALLIP, Executive Member of CCF Speech Dialogue and Hearing Special Committee, Director of Inner Mongolia Computer Society AI Special Committee, and Secretary-General of Inner Mongolia Autonomous Region Young Scientific and Technological Workers Association AI and Electronic Information Special Committee.

In real complex scenarios, human-computer interaction commonly faces problems such as modality missing and multi-speaker interference. Traditional voice interaction technologies struggle to achieve precise perception and human-like expression, severely constraining the practical application of robust dialogue intelligence. To address these challenges, this tutorial focuses on the core technical bottlenecks of human-computer voice interaction in complex scenarios, systematically presenting our innovative research in multimodal speech perception and dialogue generation, covering robust emotion recognition under missing modalities, fine-grained context graph modeling, and precise speaker identification perception technologies, effectively improving speech semantic and emotional perception stability in complex interference environments.

At the same time, we focus on human-like conversational speech synthesis technologies such as facial expression modeling, realistic emotion rendering, chain dialogue understanding, and audio-visual collaborative generation, overcoming the pain points of traditional interaction such as emotional deficiency, insufficient context adaptation, and rigid expression. Based on cutting-edge academic research, this report constructs a complete dialogue intelligence technical system of "robust perception - precise understanding - empathetic expression," analyzes development trends in the multimodal dialogue generation field, and provides new ideas and technical support for the research and development and scenario deployment of high-robustness, high-human-likeness human-computer dialogue intelligence technology.