Multilingual intelligent speech and language technologies aim at different languages, dialects, and cross-lingual application scenarios, researching how to build intelligent speech and language systems capable of understanding, generating, translating, and interacting. This is an important direction in the current development of artificial intelligence and large model technologies. With the rapid growth of global communication demands, low-resource language preservation, cross-lingual information acquisition, and multilingual human-computer interaction applications, technologies such as multilingual speech recognition, speech synthesis, speech translation, natural language processing, and cross-lingual dialogue systems have received extensive attention from academia and industry. In recent years, methods such as large-scale pre-trained models, speech and language large models, semi-supervised/self-supervised learning, cross-lingual transfer learning, and multimodal modeling have been continuously developed, providing new technical paths for solving problems such as scarce multilingual data, significant language differences, insufficient dialect and minority language resources, and limited cross-lingual generalization capabilities.
The SALT-Multilingual series of seminars has been successfully held twice, in 2023 (Nanning, CCF) and 2025 (Urumqi, NCMMSC), aiming to invite experts and scholars from academia and industry to conduct in-depth exchanges on the latest research progress, key scientific issues, system implementation methods, and industrial application scenarios of multilingual intelligent speech and language technologies. The seminar topics may cover multilingual speech recognition and generation, cross-lingual speech translation, low-resource language modeling, dialect and minority language processing, speech and language large models, multimodal speech and language understanding, multilingual human-computer dialogue, and intelligent agents. Through this third seminar, we hope to further promote academic exchanges among researchers in related fields, advance the application of multilingual intelligent speech and language technologies in education, communication, media, government affairs, cultural preservation, and international exchange, and help build a more open, inclusive, and sustainable intelligent speech and language technology ecosystem.
This report systematically introduces the recent research progress of the ASLP Laboratory at Northwestern Polytechnical University and its partners in the field of large model-driven speech recognition and understanding, covering new paradigms of speech recognition based on large language models, multi-speaker speech recognition, long audio understanding, multimodal emotion understanding, and empathetic dialogue, among other frontier directions. Meanwhile, the report will also share the team's latest achievements in open-source dataset construction, and discuss new trends in the integrated development of speech and language technologies.
This report will introduce the recent research progress of the X-LANCE Laboratory at Shanghai Jiao Tong University in the field of multilingual speech recognition and speech synthesis. With the continuous advancement of the Belt and Road Initiative, the demand for intelligent speech technologies in multilingual and multi-dialect countries in Southeast Asia, the Arab region, and others is growing daily. However, most languages still face challenges such as scarce speech data, high annotation costs, and insufficient technical coverage. How to build high-performance, low-cost, and scalable multilingual speech systems has become an important research direction.
Addressing the above issues, the report first introduces the team's exploration in multilingual speech resource construction, including the construction of large-scale multilingual datasets and evaluation benchmarks such as GigaSpeech2 and GigaSpeechBench, with a focus on sharing experience in data collection, cleaning, and quality control for low-resource languages such as Southeast Asian languages and Arabic dialects. In terms of speech recognition, the report will introduce the team's research progress based on self-supervised learning, multilingual transfer learning, and unified modeling frameworks, as well as practical results on languages such as Vietnamese, Thai, Indonesian, and Arabic. In terms of speech synthesis, the report will focus on introducing the team's recently proposed open-source models such as Habibi and X-Voice, discussing how to use unified phoneme representation, multilingual pre-training, and zero-shot voice cloning technology to achieve high-naturalness speech generation covering dozens of languages and dialects.
Improving language coverage is an important goal of multilingual speech synthesis research. This report will introduce a minimalist architecture based on the diffusion language model style, constructing a voice cloning speech synthesis model solution with extremely wide language coverage. It will also analyze solutions for speech synthesis models in different application scenarios from the perspective of model design, and discuss the future development direction of open-source speech synthesis technology.
With the development of speech large models, ASR is evolving from a traditional acoustic transcription module to a key entry point connecting multilingual speech, contextual semantics, and Agent interaction tasks. Taking NIO's self-developed speech recognition large model NIM4-ASR as an example, this report discusses how to build a parameter-efficient, low-hallucination, and personalized speech recognition large model in multilingual scenarios such as Mandarin, dialects, English, and Chinese-English mixed speech. With "acoustic fidelity, semantic alignment, and contextual controllability" as the core design principles, we introduce key technical paths such as phoneme-level modeling, multi-stage training, reinforcement learning, and hot-word personalization. Through this case, we hope to share practical thinking on the evolution of speech large models from "transcribing text" to "understanding real intent," providing reference for the research and development of multilingual intelligent speech systems and speech intelligent agents.
Zhijian Ou (Tsinghua University): ozj@tsinghua.edu.cn
Qingyang Hong (Xiamen University): qyhong@xmu.edu.cn