Young Researchers Forum

Evaluating Paralinguistic Capabilities of Speech for Natural Human-Computer Interaction: From Instruction Following to Non-verbal Vocalization
Release Time:2026/8/21 11:57:35
NCMMSC 2026 Young Researchers Forum - Liumeng Xue
Liumeng Xue
Photo
Liumeng Xue
School of Intelligence Science and Technology, Nanjing University · Tenure-track Assistant Professor

Liumeng Xue is a Tenure-track Assistant Professor at the School of Intelligence Science and Technology, Nanjing University. His main research interests include intelligent speech, speech/music/audio understanding and generation, and emotional speech generation and interaction. He received his Ph.D. from Northwestern Polytechnical University and conducted research at JD AI Lab, Tencent AI Lab, and Microsoft. He then worked as a postdoctoral researcher at The Chinese University of Hong Kong (Shenzhen) and Hong Kong University of Science and Technology.

As a co-founder, he launched a comprehensive audio generation and visualization platform, and led the development of the unified audio understanding and generation instruction dataset Audio-FLAN, which once ranked second on the Hugging Face Dataset Trending with over 100,000 downloads from top tech companies and research groups including Google, Meta, Nvidia, ByteDance, Tencent, Alibaba, Cambridge, Oxford, Tsinghua, ETH Zurich AI Center, and LAION. He participated in large speech, music, and audio generation model research including Llasa, Spark-TTS, YuE, and AudioX. His work has been published in top-tier venues such as ACL, ICLR, ICASSP, INTERSPEECH, IEEE/ACM TASLP, and Neural Networks, and has attracted attention and citations from research teams at Apple, Amazon, and other internationally renowned tech companies.

He has helped organize international conferences including ISCSLP 2026 and IEEE SLT 2024, as well as academic challenges such as MLC-SLM@INTERSPEECH 2026, SmartGlasses@SLT2026, LLM4MA Workshop@ISMIR 2025, and CoVoC@ISCSLP 2024. He regularly serves as a reviewer for ACL, ACM MM, ICASSP, INTERSPEECH, IEEE/ACM TASLP, Speech Processing Letters, and Speech Communication.

In recent years, speech generation technology has made significant progress in audio quality, naturalness, and speaker similarity. However, real human spoken communication goes far beyond "reading text aloud." The same sentence can convey completely different attitudes, emotions, and interaction intentions depending on tone, emotion, accent, speaking style, and non-verbal vocalizations such as laughter, crying, sighing, etc.

To address this issue, this talk will focus on instruction-following speech generation and non-verbal vocalization evaluation. First, we will introduce MINT-Bench, which systematically evaluates text-to-speech models' responsiveness to timbre, emotion, accent, contextual expression, compositional control, and extra-textual sound information from a multilingual instruction-following perspective. Then, we will introduce NVV-SuperBench, which further focuses on non-verbal vocalizations by constructing a Chinese-English bilingual 45-category non-verbal vocalization system, and evaluates whether models can generate specified types of laughter, crying, sighing, coughing, breathing, etc., whether they can be placed at appropriate positions, whether they are sufficiently clear and perceivable, while not degrading speech naturalness and audio quality.

The talk will combine evaluation results from multiple open-source and commercial speech generation systems to analyze the core challenges that current models still face beyond "content correctness": content consistency does not equal instruction-following capability, and overall audio quality does not equal paralinguistic control capability. Complex compositional instructions, low-SNR oral sounds, and long-duration emotional vocalizations remain obvious weaknesses of current systems. Finally, the talk will discuss how to advance non-verbal vocalization from scattered model demonstrations to comparable, reproducible, and sustainable research problems through open data, unified task settings, baseline systems, and reproducible evaluation protocols, further supporting next-generation speech generation and understanding systems for natural human-computer interaction.