Special Sessions

Special Session 10 : Speech Model Evaluation and Open-Source Ecosystem Innovation
Release Time:2026/7/4 10:01:22

Speech foundation models, general speech understanding and generation models, and various open-source speech large and small models are evolving rapidly, driving speech AI from single-task models toward multi-task, complex-scenario, and industry-grade applications. However, the field currently still faces challenges such as fragmented evaluation systems, inconsistent evaluation processes and text normalization rules, insufficient coverage of complex acoustic and linguistic scenarios, difficulty in horizontal comparison of model capabilities, and lack of unified criteria for open-source model selection and iteration. This special session focuses on two main threads: "standardized evaluation system construction" and "open-source speech model ecosystem innovation," bringing together experts from universities, open-source communities, and industry to conduct concentrated exchanges, and to distill reproducible, comparable evaluation methods, metric systems, and open-source community collaboration mechanisms that accommodate both scientific research and industrial deployment.

Kai Yu
Photo
Kai Yu
Shanghai Jiao Tong
University
Xie Chen
Photo
Xie Chen
Shanghai Jiao Tong
University
Hui Bu
Photo
Hui Bu
AISHELL
Lei Xie
Photo
Lei Xie
Northwestern
Polytechnical University
Shuai Wang
Photo
Shuai Wang
Nanjing University
Mengyao Zhu
Photo
Mengyao Zhu
Soochow University
Standardized Evaluation Framework Construction and Open Evaluation Community Co-Building Pathways
Kai Yu · Shanghai Jiao Tong University · Professor · kai.yu@sjtu.edu.cn

This talk focuses on the fragmentation of speech model evaluation, examining the difficulties in horizontal comparison, experimental reproduction, and industrial adaptation caused by inconsistent evaluation frameworks, metrics, processes, and interfaces. It discusses the core modules of a general speech evaluation framework, including process normalization, metric unification, interface standardization, and reproducible experiment protocols. Furthermore, it explores the construction model, sharing mechanisms, iteration rules, and collaborative governance schemes of an open speech evaluation community, providing an overall roadmap for co-building public evaluation infrastructure.

Speech Evaluation Standardization and Open Evaluation Ecosystem from the Research Challenge Perspective
Mengyao Zhu · Soochow University · Professor · zhu.mengyao@suda.edu.cn

This talk takes speech research challenges as the entry point, analyzing how competition rules, scoring mechanisms, task designs, and public baselines in challenges drive industry standardization. It summarizes the value of challenges in unifying evaluation baselines, standardizing model evaluation paradigms, exposing weaknesses in complex scenarios, and promoting open evaluation deployment. It also discusses how public competitions can advance the unification of evaluation rules and the normalization of evaluation processes, building a fair, open, and comparable model evaluation ecosystem.

Evaluation Standardization and Open Evaluation Deployment from the Open-Source Dataset Construction Perspective
Hui Bu · AISHELL · Founder · buhui@aishelldata.com

This talk focuses on the standardized construction and open sharing of open-source speech evaluation datasets. Addressing current issues such as single-scenario datasets, uneven quality, inconsistent annotation standards, and insufficient data in complex scenarios, it discusses dataset construction specifications, annotation standards, quality verification systems, and sharing mechanisms for general scenarios, complex acoustic scenarios, and complex linguistic scenarios. It clarifies the foundational supporting role of standardized datasets in unified evaluation, open comparison, and model robustness assessment.

Evaluation Standardization and Industrial Open Iteration from the Open-Source Model R&D Perspective
Shuai Wang · Nanjing University · Associate Professor · shuaiwang@nju.edu.cn

This talk addresses the rapid iteration of open-source speech models, analyzing—from the full chain of model R&D, fine-tuning iteration, fair comparison, and industrial deployment—how inconsistent evaluation standards constrain model capability judgment, R&D path selection, and deployment decision-making. It explores how standardized and open evaluation systems can support unified baseline evaluation, controllable training comparison, model capability grading, and differentiated capability assessment, promoting the formation of a closed-loop mechanism of "R&D—Evaluation—Iteration—Open Source—Deployment."

Kai Yu (Shanghai Jiao Tong University):kai.yu@sjtu.edu.cn

Lei Xie (Northwestern Polytechnical University):lxie@nwpu.edu.cn