Li Liu is currently an Associate Researcher and Ph.D. Supervisor at the Hong Kong University of Science and Technology (Guangzhou). She received her Ph.D. from GIPSA-lab, Université Grenoble Alpes, France. She was a postdoctoral researcher at Ryerson University, Canada. Her main research interests include speech processing, audio-visual understanding and generation, and trustworthy artificial intelligence. She has published 75 academic papers as first or corresponding author in these fields, including top-tier journals and conferences such as TPAMI, TASLP, and NeurIPS.
She currently serves as Chair of the MLSP Member Nominations & Election Subcommittee of the IEEE Machine Learning for Signal Processing Technical Committee. She served as Local Chair (China site) for ICASSP 2022. She received the Guangdong Province Young Top Talent Award and the Shenzhen Overseas High-level Talent (Peacock Talent) designation. As the principal investigator, she has hosted projects including the NSFC General Program, NSFC Key Program sub-project, NSFC Youth Program, Guangdong Provincial General Program, Guangdong Provincial Regional Joint Fund Youth Program, 2024 CCF-Tencent Rhino-Bird Project, 2023/2025 Tencent AI Lab Rhino-Bird Special Plan, 2025 CCF-Kuaishou Large Model Explorer Fund, 2022 Alibaba Innovative Research Program, and 2022/2024 Tencent Public Welfare Venture Plan.
She won the French Sephora Berribi Award for Women Scientists in Mathematics and Computer Science in 2017, the IEEE Multimedia Signal Processing Rising Star Runner-up Award, and the 2024 CCF-Tencent Rhino-Bird Project Excellence Award. Her team's paper received the Best Student Paper Nomination Award at the 16th International Conference on Social Robotics (ICSR 2024), the Outstanding Paper Award at the CVPR 2025 CV4Animals Workshop, and the Shenzhen Association for Science and Technology AI Outstanding Paper Award in 2022/2023. Her team won the Single-Model Track Champion and Agent Track Third Place in the Interspeech 2026 Audio Reasoning Challenge (156 teams from 18 countries and regions).
This talk will systematically present the team's cutting-edge breakthroughs in the fields of audio cognitive reasoning and anthropomorphic emotional speech generation. We proposed Audio-DeepThinker, which, through progressive reasoning-aware reinforcement learning, achieved for the first time the autonomous emergence of structured chain-of-thought in large audio language models, topping multiple deep reasoning benchmarks.
On this basis, we constructed PhyAVBench, the world's first audio-visual generation evaluation platform for physical commonsense, and innovated a comparative physical response scoring mechanism, pushing generative models from "audio-visual synchronization" to a new height of "physical interpretability." In the direction of emotional generation, we developed EmoSteer-TTS, which for the first time achieves training-free, activation-level fine-grained emotional control in speech synthesis, opening up a new paradigm for controllable human-computer interaction.
Furthermore, the team launched AudioGenie, a unified multi-agent collaborative framework that bridges the generation closed loop from text, images, and video to various types of audio including speech, sound effects, music, and songs. The above works collectively build a complete technical system from audio logical reasoning to emotional intelligent generation, laying key theoretical and application foundations for next-generation embodied intelligence, multimodal interaction, and physical world simulation.