Special Sessions

Special Session 1 : ChinaVoices Challenge 2026 Chinese Multi-Dialect Speech Recognition Challenge
Release Time:2026/7/2 22:46:31

In recent years, automatic speech recognition (ASR) technology has been widely deployed in intelligent voice assistants, meeting transcription, intelligent customer service, and voice interaction systems. However, in real-world scenarios, user speech is not limited to standard Mandarin but encompasses a broad spectrum of accents and dialects. Consequently, multi-dialect speech recognition in complex acoustic environments has garnered increasing attention from both academia and industry.

Chinese dialects are characterized by extensive geographic distribution, significant phonetic divergence, diverse lexical expressions, high annotation costs, and unbalanced resource development. Compared with Mandarin speech data, dialect speech data generally suffer from scarcity of resources, insufficient coverage, and limited annotation scale, constituting a typical low-resource ASR scenario. While existing ASR models perform well on standard Mandarin, they still face substantial challenges in multi-dialect, low-resource, and cross-regional generalization settings, struggling to robustly adapt to dialectal speech inputs in practical applications.

Meanwhile, the field of Chinese dialect speech processing still lacks a unified, publicly available, and reproducible standardized evaluation platform. Existing studies vary considerably in data sources, dialect categories, test-set partitioning, evaluation metrics, and experimental configurations, making fair, systematic, and reproducible performance comparisons difficult and, to some extent, hindering the sustained advancement of Chinese dialect ASR technology.

Therefore, leveraging NCMMSC 2026 as a platform, this challenge aims to establish a unified evaluation dataset, standardized metrics, and competitive procedures for Chinese multi-dialect speech recognition and dialect identification, providing an open, fair, and reproducible benchmark for both academia and industry. The challenge will release manually transcribed and quality-controlled speech data covering more than ten Chinese dialects and regional accents to support low-resource dialect modeling and cross-regional generalization evaluation. Each dialect dataset has undergone manual transcription and quality assurance, offering reliable support for low-resource dialect modeling and cross-dialect generalization assessment. The challenge comprises two tracks: Chinese multi-dialect identification and Chinese multi-dialect speech recognition, evaluating models' capability to discriminate dialect categories and to transcribe multi-dialect speech content, respectively. Through this challenge, we seek to further promote the development of low-resource Chinese dialect ASR technology and facilitate the construction of Chinese multi-dialect speech resources and the refinement of standardized evaluation systems.

Lei Xie
Photo
Lei Xie
Northwestern Polytechnical University
Liumeng Xue
Photo
Liumeng Xue
Nanjing University
Hexin Liu
Photo
Hexin Liu
Nanyang Technological University
Xian Shi
Photo
Xian Shi
Alibaba Tongyi Laboratory
Jie Hu
Photo
Jie Hu
Beijing Huiting Technology Co., Ltd.
Intelligent Speech Systems for Southern Min (Hokkien)
Qingyang Hong · Xiamen University · Professor · qyhong@xmu.edu.cn

This talk first introduces the application background of Southern Min (Hokkien), followed by the romanization scheme designed by the Xiamen University team. The presentation then focuses on the technical approaches adopted for Southern Min speech recognition, including early hybrid architectures and end-to-end models. By employing progressive training strategies and semi-supervised learning mechanisms, the latest version achieves a substantial improvement in character recognition accuracy while simultaneously supporting Mandarin recognition. Unlike other commercial systems, this system displays all recognition results in Mandarin characters, offering greater user-friendliness. The talk will also introduce the Southern Min speech synthesis system designed for different local accents (e.g., Xiamen accent and Quanzhou accent). A live system demonstration will be provided.

Multilingual Speech Data Resource Construction for Southeast Asia and Belt-and-Road Countries
Xie Chen · Shanghai Jiao Tong University; Shanghai Innovation College · Associate Professor · chenxie95@sjtu.edu.cn

With the advancement of the Belt and Road Initiative, the demand for multilingual intelligent speech technology in Southeast Asia, the Middle East, and other regions continues to grow. However, many languages and dialects still face challenges such as scarce speech resources, lack of evaluation benchmarks, and insufficient foundational model capabilities. This talk will present our recent explorations and progress in multilingual speech resource construction and foundational model research.

The presentation first introduces the GigaSpeech2 large-scale multilingual speech dataset constructed by our team, covering multiple low-resource languages in Southeast Asia and providing high-quality training resources for multilingual speech recognition and synthesis research. Next, we introduce the GigaSpeechBench multilingual speech evaluation benchmark, which systematically assesses the performance of current open-source and commercial models on low-resource languages, dialects, and cross-lingual scenarios, and analyzes the primary technical challenges. Furthermore, the talk will share our research progress on Arabic and Arabic dialects, including large-scale speech data construction, multi-dialect speech recognition systems, and multi-dialect speech synthesis models.

Through these efforts, we aim to provide high-quality data resources, evaluation benchmarks, and foundational model support for the development of multilingual speech technology in Belt-and-Road countries, and to discuss future directions for multilingual speech intelligence.

G-STAR: An End-to-End Framework for Long-Form Multi-Speaker Speech Recognition
Shuai Wang · Nanjing University · Associate Professor · shuaiwang@nju.edu.cn

This talk focuses on the speaker-attributed speech recognition task in long-form multi-party meeting scenarios, i.e., identifying "who spoke what and when." Addressing the issues of speaker identity inconsistency and imprecise temporal boundaries in existing chunk-based inference methods, we introduce the end-to-end framework G-STAR. This method integrates speaker tracking caches with Speech-LLM, maintaining speaker identities across chunks and incorporating speaker cues into the large language model decoding process to achieve timestamped and globally speaker-labeled speech transcription. Experimental results demonstrate that G-STAR outperforms traditional cascaded approaches and existing end-to-end models on multiple meeting speech datasets, exhibiting strong capabilities in long-form multi-speaker speech understanding.

Dialect Data Construction and Speech Technology Empowerment
Jie Hu · Beijing Huiting Technology Co., Ltd. · Director · hujie@huitingtech.com

Behind the surging demand for dialect technologies, the core bottlenecks remain the heterogeneity of accents, large phonetic divergence, lack of standardization, and data scarcity. Major technology companies commonly face insufficient data for low-incidence dialects, weak cross-domain generalization, and low accuracy in noisy environments, creating an urgent need for standardized, comprehensive, and high-quality dialect databases. Huiting has established a full-process standardized production system, completing territory-wide collection tailored to regional dialect characteristics and preserving original acoustic features in professional recording environments. Relying on unified annotation protocols and multi-level quality inspection, we produce high-consistency, premium-quality dialect corpora to empower speech technology research and development.