Zheng Lian (IEEE/CCF Senior Member) is currently an Associate Professor and Ph.D. Supervisor at Tongji University. In recent years, he has conducted a series of research works in affective computing, human-computer interaction, and large models, with over 100 published papers and patents, including international journals such as TPAMI, TNNLS, TASLP, TAFFC, and international conferences such as ICML, NeurIPS, and ICLR. He has accumulated over 5,400 citations on Google Scholar (H-index: 38) and was listed among the world's top 2% scientists. As one of the principal contributors to "Key Technologies and Systems for Multimodal High-Robustness Fine-Grained Affective Analysis," he received the First Prize for Technical Invention from the Chinese Institute of Electronics. He released the multimodal emotion recognition database MER and organized challenges and workshops around this database for four consecutive years at ACM Multimedia and IJCAI. He serves as Associate Editor for IEEE TAFFC/IEEE TASLP/PR, Area Editor for Information Fusion, Area Chair for ACM Multimedia and ACL ARR, and Dataset Co-Chair for ACM Multimedia 2026.
Emotion is closely linked to cognition, decision-making, and behavior, playing a key role in the field of speech interaction. Emotion representation methods aim to map complex human emotions into quantifiable values. Currently, there are two main paradigms in emotion representation: categorical emotion representation and dimensional emotion representation. Categorical emotion representation is based on psychological theories that divide human emotions into discrete labels. However, human emotions are far more complex than six basic emotions, and restricting the emotion space to basic categories inevitably overlooks some subtle emotions. Unlike categorical emotions, dimensional emotion representation models human emotions as a point in a continuous multi-dimensional space, capable of modeling fine-grained emotions. However, dimensional representation is more abstract than categorical representation and inconsistent with human intuitive perception of emotions, limiting its application in downstream tasks.
This talk will summarize some of our recent attempts in this field, hoping to leverage the rich vocabulary and multimodal perception capabilities of multimodal large models to transition from discriminative emotion recognition to fine-grained, interpretable generative emotion understanding, including EMER, OV-MER (ICML25), AffectGPT (ICML25), EmoPrefer (ICLR26), and AffectGPT-RL, as well as the exploration of affective computing in embodied intelligence, RobotEQ, hoping to provide some references for the affective computing community.