Special Sessions

Special Session 8 : Empathic Computing and Cognitive Interaction
Release Time:2026/7/4 10:01:05

The core purpose of the topic "Computational Empathy and Cognitive Interaction" is to break the cold "command-response" paradigm of traditional Human-Machine Interaction (HMI). By endowing machines with the ability to recognize, understand, and even simulate human emotions and cognitive states, we aim to build a new generation of interactive systems possessing Social Intelligence. This endeavor is not merely about enhancing technical usability, but also about addressing the urgent need for ethical and human-centric development in Artificial Intelligence.

This special session focuses on two core directions, "Computational Empathy" and "Cognitive Interaction," covering cutting-edge topics including emotion recognition in speech, missing-modality emotion recognition, multimodal emotion and intent joint recognition, reasoning-augmented multimodal empathetic response generation, and empathy modeling in active inference for robotics. We bring together experts from universities, research institutes, and industry to explore the leap from signal processing to semantic and emotional understanding, and to promote deep applications of human-computer interaction in mental health, education, healthcare, and other scenarios.

Dongyan Huang
Photo
Dongyan Huang
Shenzhen Joylift AI
Co., Ltd.
Ziping Zhao
Photo
Ziping Zhao
Tianjin Normal University
Zixing Zhang
Photo
Zixing Zhang
Hunan University
Ya Li
Photo
Ya Li
Beijing University of
Posts and Telecommunications
Bin Liu
Photo
Bin Liu
Institute of Automation,
Chinese Academy of Sciences
Investigation of Empathy Modeling in Active Inference for Robotics
Dongyan Huang · Shenzhen Joylift AI Co., Ltd. · Chief Scientist · dongyan_huang@xinyang-ai.com

To elucidate the emergence of embodied empathy in physical agents, we propose a computational framework grounded in Active Inference. Utilizing explicit perspective-taking and self-other model transformation, this framework enables multi-agent interaction from a single generative model without maintaining separate models for each interlocutor; the agent dynamically transforms its self-model to infer others' beliefs, goals, and action tendencies.

We instantiate this framework within a multi-agent interactive resource allocation task. Results demonstrate that empathy-driven cooperation emerges robustly only under reciprocal conditions, whereas asymmetric empathy leads to systematic exploitation. Extending the framework to learning-enabled agents via Bayesian model updating, we observe that while open models converge rapidly, long-term cooperative stability remains governed by intrinsic empathy parameters.

These findings position Active Inference as a foundational principle for socially aligned AI—facilitating coordination through internal simulation rather than behavioral mimicry—and offer a new paradigm for designing trustworthy, interpretable, and ethically grounded empathic agents.

Emotion Recognition in Speech via Learning on Self-Estimated Labels Using a Dual-Branch MoE-Based Network
Ziping Zhao · Tianjin Normal University · Professor · zhaoziping@tjnu.edu.cn

Speech Emotion Recognition (SER) aims to recognize emotions from speech signals. Existing studies typically use majority-vote annotations as ground-truth labels for model training, but overlook the discrepancy between ground-truth labels and self-estimated labels caused by emotional ambiguity and subjective human perception. To address this issue, this talk proposes a speech emotion recognition method based on a dual-branch Mixture-of-Experts (MoE) network that learns from self-estimated labels.

The method comprises a speech representation extraction module and a dual-branch MoE module, which alleviates the label discrepancy by incorporating self-estimated labels into the supervision process. Experimental results on multiple emotional speech datasets demonstrate that the proposed method outperforms existing state-of-the-art approaches, validating its effectiveness for speech emotion recognition tasks.

Mitigating Modal Semantic Inconsistency via Uncertainty-Aware Soft Labels for Missing-Modality Emotion Recognition
Zixing Zhang · Hunan University · Professor · zixingzhang@hnu.edu.cn

In multimodal emotion recognition under missing-modality conditions, adapting to various missing scenarios is crucial. Existing methods mainly focus on feature reconstruction, but largely ignore the inconsistency of emotional expression across different modalities. Overlooking this inconsistency is risky because available modalities may convey misleading information, causing the model to learn incorrect mapping relationships.

To address this neglected issue, this talk proposes the F2M-Net framework, which explicitly considers and mitigates these conflicts. Unlike previous methods that ignore modal inconsistency, we utilize an adaptive soft-label learning mechanism to dynamically calibrate supervision signals based on instance-level input reliability. This mechanism reduces the risk of overfitting to misleading cues, thereby improving robustness. Experimental results show that F2M-Net outperforms existing baseline methods on multiple datasets.

Swapped-Query Cascaded Fusion for Multimodal Emotion and Intent Joint Recognition
Ya Li · Beijing University of Posts and Telecommunications · Professor · yli01@bupt.edu.cn

Emotion recognition and intent recognition are key tasks in natural language processing and multimodal computing, and also the cornerstone of intelligent human-computer interaction. However, traditional research often treats them as independent downstream tasks, ignoring their intrinsic causal relationship. Consequently, when handling complex contexts such as sarcasm, existing models often suffer from recognition bias due to insufficient cross-task information interaction.

To address these challenges, this talk proposes an improved framework based on the multimodal joint recognition baseline model EI2. First, to solve the reproducibility issues of the baseline model, a robust multimodal feature extraction pipeline is designed to achieve cross-modal dimensional alignment. Second, to address the limitations of static query vectors in the cascaded fusion stage, a novel Swapped-Query cascaded fusion mechanism is proposed, which reconstructs the underlying attention retrieval logic to force the emotion branch and intent branch to leverage each other's features, thereby promoting deep cross-task interaction. In addition, to alleviate performance constraints caused by long-tail data distribution, a hybrid loss function combining cross-entropy and Focal Loss is introduced. Finally, to enhance model interpretability, an attention visualization module is incorporated to analyze cross-modal dynamic attention migration.

Experimental results on the MC-EIU multimodal dataset show that the proposed method achieves significant performance improvements: the average accuracy of intent recognition increases from 46.60% to 48.25%, and the Unweighted Average Recall (UAR) for minority classes significantly improves from 25.18% to 31.43%; the accuracy of emotion recognition simultaneously increases by 3.12%. Case studies further confirm that the model effectively captures complementary relationships between modalities at the algorithmic level, demonstrating its dual capabilities in precise recognition and deep semantic understanding. This talk provides an effective and interpretable framework for multimodal emotion and intent joint understanding in complex conversational scenarios.

RA-MERG: A Reasoning-Augmented Multi-modal Empathetic Response Generation Benchmark
Bin Liu · Institute of Automation, Chinese Academy of Sciences · Associate Professor · liubin@nlpr.ia.ac.cn

Empathetic dialogue systems aim to replicate human-like interaction by perceiving and expressing multimodal emotions. However, current research is limited by the scarcity of high-quality multimodal datasets and the dominance of cascaded generation pipelines, leading to misalignment between linguistic content and paralinguistic expression.

This talk addresses these challenges by constructing the RAMERG (Reasoning-Augmented Multi-modal Empathetic Response Generation) dataset, which contains rich high-level reasoning annotations. Furthermore, we propose a novel end-to-end speech and text generation framework. Unlike traditional pipeline methods, this framework jointly models semantic and acoustic representations to directly generate coherent text and speech waveforms, thereby maintaining emotional consistency from input perception to output generation. Extensive experiments demonstrate that the proposed framework exhibits superior empathetic capabilities compared to baseline models while ensuring high-quality text and audio generation.

This talk lays a solid foundation for future research on multimodal empathetic response generation.

Dongyan Huang (Shenzhen Joylift AI Co., Ltd.):dongyan_huang@xinyang-ai.com

Ziping Zhao (Tianjin Normal University):zhaoziping@tjnu.edu.cn