Xiaoman Wang is a Senior Engineer at the Data Platform Center, Shanghai Artificial Intelligence Laboratory. She has extensive experience in innovative corpus construction and has long been engaged in multilingual and multimodal corpus resource construction and data governance. She has participated in the drafting of multiple national standards and team standards related to corpora. Her research interests mainly include multimodal corpus construction, low-resource language data construction, intelligent data processing, and high-quality annotation systems.
This talk will focus on the language interoperability needs of countries along the Belt and Road, ASEAN, and BRICS, introducing the construction ideas, key technologies, and practical achievements of the "Wanjuan · Silk Road" multimodal low-resource language corpus. Currently, low-resource language data generally faces challenges of insufficient quantity and suboptimal quality. The talk will introduce how to innovatively build high-quality multimodal multilingual corpora including audio, video, and text, how to achieve industrial-grade high-quality corpus construction through a collaborative closed-loop mechanism of "country-specific experts + intelligent processing + low-resource language experts", and the practical experience in real-world application scenarios such as enterprise globalization. It aims to provide strong support for eliminating "language islands", promoting mutual learning among civilizations, and building high-quality digital bridges.