Young Researchers Forum

Research on Low-Bitrate Neural Speech Codec Methods
Release Time:2026/8/21 11:56:05
NCMMSC 2026 Young Researchers Forum - Yang Ai
Yang Ai
Photo
Yang Ai
National Engineering Research Center of Speech and Language Information Processing, USTC · Associate Researcher

Yang Ai is an Associate Researcher at the National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China (USTC). His main research interests include speech coding, speech enhancement, speech synthesis, and audio quality assessment. He has published over 70 papers in renowned journals and conferences in the speech technology field, with nearly 60 papers as first or corresponding author. He currently hosts projects funded by the National Natural Science Foundation of China and the Anhui Provincial Natural Science Foundation, and participates in multiple projects including the Strategic Priority Research Program and the National Key R&D Program.

In terms of awards, he was selected as a "Xiaomi Young Scholar" in 2024, won the Vocoder Track Champion (first completer) of the Interspeech 2024 Discrete Speech Challenge, and received the Best Paper Award (corresponding author) of the 18th National Conference on Man-Machine Speech Communication (NCMMSC).

In recent years, neural speech codecs have demonstrated significant application value in low-bitrate speech compression, real-time voice communication, and on-device speech storage. However, maintaining high speech reconstruction quality at even lower bitrates remains a challenging problem. This talk will introduce our research on low-bitrate neural speech coding methods, focusing on two aspects: overall codec framework design and quantization strategy improvement.

First, we will introduce FMelCodec, an ultra-low-bitrate neural speech codec proposed by our team. This method uses mel-spectrogram as the modeling target and builds an "encode-refine-reconstruct" framework. Through single-codebook discrete encoding, conditional flow matching spectral refinement, and vocoder reconstruction, it achieves high-quality speech reconstruction at bitrates as low as 250 bps.

Second, we will introduce two low-bitrate neural speech coding methods focusing on quantization strategy improvements. One is VoCodec for streaming voice communication scenarios, which adaptively allocates quantization resources according to the different perceptual importance of voiced and unvoiced sounds, thereby reducing the overall bitrate while maintaining speech quality. The other is AKDQ for residual vector quantization neural speech codecs, which leverages temporal redundancy between adjacent speech frames by partitioning frames into key frames and differential frames, and achieves more efficient adaptive quantization through a causal bitrate scheduling mechanism.