Young Researchers Forum

Towards More Stable Fully Continuous Autoregressive Speech Synthesis
Release Time:2026/8/21 11:59:49
NCMMSC 2026 Young Researchers Forum - Da Zheng, Hankun Wang, Bohan Li
Da Zheng
Photo
Da Zheng
Xiaohongshu (RED) dots Speech Team
Hankun Wang
Photo
Hankun Wang
Shanghai Jiao Tong University X-LANCE · Ph.D. Student
Bohan Li
Photo
Bohan Li
Shanghai Jiao Tong University · Ph.D. Student

Da Zheng received his Master's degree from Shanghai Jiao Tong University, supervised by Prof. Kai Yu. He built the industry-leading large-scale Chinese speech recognition system yitu-asr. He is currently working on text and multimodal model research and development at the Xiaohongshu dots speech team, and is conducting research in the direction of full-duplex speech.

Hankun Wang is a Ph.D. student at the Cross-media Language Intelligence Lab (X-LANCE), Shanghai Jiao Tong University, supervised by Prof. Kai Yu. He received his bachelor's degree from Harbin Institute of Technology. He is currently interning at the Xiaohongshu dots speech team. His main research direction is speech representation and synthesis. He has published more than 10 papers in flagship conferences and journals such as ICASSP, Interspeech, SLT, and AAAI, and serves as a reviewer for conferences such as ICASSP and Interspeech.

Bohan Li received his B.Eng. degree in Computer Science and Technology (ACM Class) from Shanghai Jiao Tong University in 2024, and is currently pursuing his Ph.D. at Shanghai Jiao Tong University, supervised by Prof. Kai Yu. He is currently interning at the Xiaohongshu dots speech team. His main research interests are speech synthesis and understanding. He has published papers in multiple top-tier conferences in the fields of speech and language processing and machine learning, including ICASSP, Interspeech, NeurIPS, and AAAI.

This talk mainly introduces the speech exploration progress of the Xiaohongshu dots team. It is divided into three parts: continuous representation autoregressive speech synthesis dots.tts, its dependent generation-understanding unified speech representation HoliTok, and the unified fine-grained speech editing model dots.tts.edit built on top of dots.tts.

dots.tts is an open-source 2B-parameter fully continuous end-to-end autoregressive TTS foundation model. It utilizes semantic AudioVAE (HoliTok), preserving continuous speech details while maintaining model generation stability. On the Seed-TTS-Eval Chinese-English zero-shot voice cloning benchmark, dots.tts achieved the best average results. CFG-aware MeanFlow distills the 16-step teacher trajectory with CFG into interval average velocity, requiring only 2 to 4 single-branch function evaluations during inference. dots.tts also supports a 1T1A dual-stream mode, providing low-latency speech output capability for full-duplex duplex mode.

HoliTok encodes 48 kHz speech into 25 Hz, 128-dimensional continuous representations. The model adopts a progressive training strategy, maintaining signal-level reconstruction fidelity while preserving good learnability of the representation space and introducing semantic information. We built a unified AR+DiT model for simultaneous downstream modeling of speech synthesis and speech recognition. The same representation sequence supports both generation-oriented tasks and unified generation-understanding downstream tasks.

dots.tts.edit uses the continuous autoregressive dots.tts as the base model. Taking source speech and structured instructions anchored to transcribed text as input, it directly locates text segments or pause boundaries that need editing, uniformly completing text, emotion, prosody, pause, and compositional editing. Experiments show that without complex semantic encoders or other additional modules, and without substantially modifying the backbone of the base model, a continuous autoregressive TTS foundation model can be transformed into a model suitable for unified fine-grained speech editing.