Special Sessions

Accessible Interaction and Pathological Speech Processing by Graduate Students with Disabilities
Release Time:2026/7/4 10:00:37

This special session focuses on frontier directions such as accessible human-computer interaction, pathological speech processing, intelligent assisted communication, multimodal large models, and information accessibility. It invites graduate student speakers with disability experience to share their scientific research, technical exploration, and practical reflections. The session addresses the real needs of people with disabilities in learning, communication, research, employment, and daily life, and discusses topics such as dysarthric speech recognition, intelligent voice-assisted communication, accessible interaction design, multimodal content generation, and academic information accessibility, exploring the application of artificial intelligence technology in communication accessibility, information accessibility, and educational accessibility.

Unlike general technology-for-disability topics, this session highlights the agency of graduate students with disabilities in accessible intelligent technology research. Graduate students with disabilities are not only important users of accessible technology but also direct observers, experiencers, and researchers of relevant issues. Drawing on their own experiences and professional training, they are able to provide more authentic problem definitions, more nuanced user understanding, and clearer application orientations for accessible intelligent technology, helping to advance AI technology from "designing for people with disabilities" to "designing with people with disabilities."

This session will promote cross-disciplinary integration among artificial intelligence, speech processing, human-computer interaction, special education, rehabilitation assistance, and information accessibility, and encourage researchers to carry out technological innovation addressing communication barriers, interaction barriers, and information access barriers in real-world scenarios. Through academic presentations and exchanges by graduate students with disabilities, this session not only showcases the spirit of self-improvement and serving society through research among young people with disabilities, but also helps promote disability-assistive technology achievements to be more user-centered, better adapted to real-world scenarios, and more accessible and inclusive.

The establishment of this session will provide a high-level academic platform for graduate students with disabilities, foster in-depth dialogue among academia, industry, educational institutions, rehabilitation organizations, and people with disabilities, and promote technologies such as dysarthric speech recognition, intelligent assisted communication, multimodal interaction, and accessible content generation from the laboratory to real-world application scenarios, providing technical and talent support for building an "inclusive and barrier-free" intelligent society.

Sujing Wang
Photo
Sujing Wang
Institute of Psychology,
Chinese Academy of Sciences
Xixin Wu
Photo
Xixin Wu
The Chinese University of Hong Kong
Mingming Fan
Photo
Mingming Fan
HKUST (Guangzhou)
Nan Yan
Photo
Nan Yan
SIAT,
Chinese Academy of Sciences
Shiqi Yu
Photo
Shiqi Yu
Southern University of Science and Technology
Yanling Li
Photo
Yanling Li
Xinyang Normal University
Ying Chen
Photo
Ying Chen
Nanjing University of Science and Technology
Jian Zhao
Photo
Jian Zhao
Changchun University
Dengfeng Yao
Photo
Dengfeng Yao
Beijing Union University
Wen Hu
Photo
Wen Hu
Fujian Normal University
Shucheng Huang
Photo
Shucheng Huang
Jiangsu University of Science and Technology
Research on Text Input Assistance Technology for Cerebral Palsy Patients
Yunfei Bi · Xinyang Normal University · Master's Student · byf19990715@163.com

This presentation focuses on the inefficient text input dilemma and digital communication barriers faced by people with cerebral palsy in human-computer collaboration, deeply analyzes their hand-eye motor coordination strategies during text input, and demonstrates the AR intelligent interactive keyboard independently developed by the team based on this research, aiming to achieve immersive and efficient input through visual reconstruction. It also prospectively discusses virtual keyboard interaction strategies and accessible text input design paradigms for cerebral palsy patients.

Can Multimodal LLM Models Help Improve Dysarthric Speech Recognition?
Bai Liu · School of Computer Science and Technology, Changchun University · Master's Student · 1169656535@qq.com

The speech recognition accuracy of dysarthric patients is far lower than that of normal populations. This study explores whether multimodal LLMs can improve this issue—by fusing acoustic signals with the contextual understanding capabilities of large language models, and fully utilizing cross-modal complementary information to compensate for the acoustic deficiencies of dysarthric speech.

Structured Contract-Driven Automatic Generation of Academic Paper Presentation Videos
Jingqi Wu · Southern University of Science and Technology, Zhongguancun Academy · PhD Student · jingqiwu2020@163.com

Automatically converting academic papers into presentation videos requires synchronously generating slides, speech, subtitles, cursor animations, and presenter avatars across multiple modalities while ensuring semantic consistency among them. Existing methods treat each modality as an independent generation task, using the original paper text as the respective input source for separate inference, leading to cross-modal semantic drift, and any local modification requires re-running the entire pipeline.

This presentation introduces PaperContract, a structured contract-driven academic video generation framework. Its core idea is to introduce a persistent structured contract between the paper and downstream modalities—a unified specification for the narrative text, visible content, and visual assets of each slide. All modules use this contract as the sole information source, fundamentally eliminating redundant semantic inference. During generation, slides are produced through a QA-guided verification-refinement loop; speech is synthesized only once per slide, and word-level timestamps are obtained through forced alignment to build a shared temporal contract that synchronously drives subtitles, cursor positioning, and Talking Head rendering, ensuring cross-modal synchronization by design rather than post-hoc alignment. During interactive modification, users issue modification instructions in natural language, and the system only updates the affected units in the contract and selectively invalidates corresponding caches, while unaffected slides, audio, and animations are directly reused, significantly reducing modification costs.

Using Intelligent Speech Technology to Help People with Dysarthria Communicate with Others
Xinran Zhao · School of Computer Science, Jiangsu University of Science and Technology · Master's Student · 3212835198@qq.com

Dysarthria is a motor speech disorder caused by neurological damage. Due to the acoustic differences from standard speech and the scarcity of data, directly using general speech recognition models often leads to performance degradation. Combining existing research, we mainly investigate model architecture adaptation, acoustic feature analysis, curriculum learning training strategies, and other directions to improve dysarthric speech recognition performance.

Evaluation of Information Accessibility Structured Caption Generation for Hearing-Impaired Groups
Jie Shi · School of Robotics, Beijing Union University · Master's Student · 1807872045@qq.com

With the deepening of information accessibility construction in China, captions have become essential infrastructure for hearing-impaired groups and elderly people to access audio and video information. However, in real communication scenarios with high immediacy and information density, traditional speech recognition technology often faces challenges such as long social latency, easy loss of core information, and lack of readability in layout. This paper introduces the "Information Accessibility Structured Caption Generation Evaluation for Hearing-Impaired Groups" task at CCL 2026 (CCL26-Eval). This evaluation focuses on the complete technical pipeline from audio and video input to highly usable structured caption generation. To approximate real-world applications, the task establishes a "PC track" for exploring performance limits and a "mobile track" under constrained resources, and subdivides into two subtasks: basic caption generation and structured readable caption generation. To address the high complexity of spoken expression in real corpora and noise in core information annotation, this evaluation constructs a hybrid evaluation system of "automatic metrics with manual verification and arbitration." This evaluation aims to promote AI caption technology that comprehensively considers transcription accuracy, response speed, and information structured presentation, providing a unified benchmark and high-quality multi-scenario datasets for model research and development and edge deployment in the information accessibility field.

A Comparative Study of Chinese Science and Technology Image Construction in DeepSeek-Themed English and Chinese News Discourse from the Perspective of Critical Metaphor
Qinghe Huang · Fujian Normal University · Master's Student · 2945370366@qq.com

As international technological competition increasingly focuses on artificial intelligence, China's independently developed large model DeepSeek has attracted international media attention, providing a key case for examining the construction and reconstruction of China's science and technology image within the existing international discourse landscape. Using critical metaphor analysis as the theoretical framework, this study builds a self-constructed DeepSeek-themed English and Chinese news corpus (January 2025 to May 2026, containing approximately 150 reports each from People's Daily and The New York Times), and employs the MIPVU metaphor identification procedure, corpus quantitative statistics, and critical discourse qualitative analysis to compare how Chinese and American media construct China's science and technology image through conceptual metaphors. Different metaphor choices influence audience cognition and judgment of China's technological development through mapping differences and framing effects, thereby participating in the shaping of national science and technology image. The study aims to answer three core questions: What core conceptual metaphors are used by both sides? How do these metaphors frame China's technological development and reveal what value orientations? What kind of technological discourse power and geopolitical games do the systematic differences in metaphor choices reflect? It is expected to clarify the metaphor distribution maps of both countries, reveal the cognitive models and value differences behind metaphor differences, and provide specific discourse strategies for constructing Chinese science and technology narratives with cultural autonomy and international persuasiveness. This study is both an empirical deepening of critical metaphor analysis in the field of artificial intelligence discourse and provides theoretical support and practical insights for understanding and optimizing the global communication mechanism of China's science and technology image.

Research on Intelligent Dysarthric Speech Recognition Methods for Accessible Interaction
Xinchen Kang · School of Computer Science and Engineering, Nanjing University of Science and Technology · PhD Student · kxc4088@163.com

Dysarthric speech recognition is an important issue in accessible human-computer interaction and intelligent rehabilitation assistance. Due to unclear pronunciation, decreased speech intelligibility, and significant individual differences, existing automatic speech recognition models still have high error rates in pathological speech scenarios. For fixed-vocabulary dysarthric speech recognition tasks, this study uses self-supervised speech models as the foundation, constructs a CTC and word classification dual-head learning framework, and introduces a clarity-aware acoustic-edit distance vocabulary rescoring method to enhance the model's ability to correct approximate words and substitution errors. Experiments are conducted on the UASpeech dysarthric speech dataset, and evaluation is performed from the perspectives of overall recognition rate, speaker level, and different clarity levels. Results show that combining word-level supervision and clarity-aware decoding strategies helps improve the robustness of pathological speech recognition. This research can provide technical support for accessible communication, intelligent voice interaction, and rehabilitation assessment for people with disabilities.

Sujing Wang (Institute of Psychology, Chinese Academy of Sciences): wangsujing@psych.ac.cn

Xixin Wu (The Chinese University of Hong Kong): wuxx@se.cuhk.edu.hk

Mingming Fan (HKUST (Guangzhou)): mingmingfan@ust.hk

Nan Yan (SIAT, Chinese Academy of Sciences): nan.yan@siat.ac.cn

Shiqi Yu (Southern University of Science and Technology): yusq@sustech.edu.cn

Yanling Li (Xinyang Normal University): lyl75@163.com

Ying Chen (Nanjing University of Science and Technology): ychen@njust.edu.cn

Jian Zhao (Changchun University): zhaojian@ccu.edu.cn

Wen Hu (Fujian Normal University): wendyhu6@163.com

Dengfeng Yao (Beijing Union University): tjtdengfeng@buu.edu.cn

Shucheng Huang (Jiangsu University of Science and Technology): schuang@just.edu.cn