In recent years, with the emergence and rapid iteration of generative AI technology, the naturalness and speaker similarity of generated audio have been significantly improved, bringing issues such as speech fake detection, speaker recognition and verification, and training sample tracing. On the one hand, existing audio information processing technologies still face challenges in stability, robustness, generalization, and interpretability; on the other hand, new problems such as privacy leakage and voiceprint spoofing urgently need to be solved. Trustworthy audio processing technology aims to address algorithmic robustness, privacy, and stability issues in audio information processing. This special session invites young researchers who have made achievements in the field of trustworthy audio in recent years to showcase this emerging topic of trustworthy audio processing technology, discuss privacy protection and data tracing issues in existing audio information processing technologies, and target the technology, evaluation, and application scenarios of trustworthy audio information processing under the new business format of generative audio processing technology. The session gathers forces from academia and industry to showcase cutting-edge trustworthy audio processing technologies such as speech fake detection, voiceprint protection and recognition, and training data tracing, demonstrate the application bottlenecks of related technologies in open scenarios, and lay a foundation for compliant application of next-generation audio security technology based on trustworthy artificial intelligence.
Speaker representation learning aims to extract compact and discriminative embedding representations to capture speaker-specific acoustic features while minimizing the impact of linguistic content and acoustic environment variations. However, existing methods still face several challenges: first, it is difficult to effectively reduce the representation differences of the same speaker under different speech conditions; second, the efficient adaptation of large-scale pre-trained speech models to speaker verification tasks still involves high computational and storage costs; third, speaker identity information and speech content information are often coupled with each other, making it difficult to directly obtain content-independent speaker representations.
This presentation will introduce three research works addressing the above issues. First, we propose a supervised contrastive learning framework combined with additive angular margin to improve intra-class compactness and inter-class separability. By maximizing the mutual information between frame-level features and speaker representations, this method can retain key speaker-related information under various data augmentation conditions. Second, we explore parameter-efficient fine-tuning methods for pre-trained Transformer speech models, including dynamic prompt tuning and spectrum-aware LoRA, achieving effective adaptation to speaker verification tasks while significantly reducing computational and storage overhead. Finally, we propose a diffusion model-based variational framework to disentangle speaker timbre from speech content, thereby obtaining content-independent speaker representations that are more robust to linguistic content variations. Experimental results on the CN-Celeb, VoxCeleb, and CU-MARVEL datasets show that the above methods can effectively improve the robustness, transferability, and interpretability of speaker representations. This study provides new insights for more trustworthy speaker representation learning by enhancing discriminative ability, adaptation efficiency, and content invariance in speaker verification systems.
Deepfake speech detection is one of the key tasks in trustworthy audio processing. With the rapid development of current generative speech technology, forged speech has been continuously improving in terms of naturalness and speaker similarity. Meanwhile, the rapid iteration of generation models and the complexity of cross-lingual and cross-scenario detection conditions have also made the stable application of existing methods in real-world open scenarios challenging. This presentation will systematically review the key issues and research progress in deepfake speech detection, and introduce a series of studies conducted by the team focusing on the fine-grained differences between authentic speech and deepfake speech, covering detection scenarios such as global forgery and local forgery, including background noise and low-frequency residual information modeling, multi-subband dynamic representation, and temporal difference modeling. This presentation will summarize the team's explorations and achievements in detection generalization and interpretability, and discuss new paradigms for trustworthy deepfake speech detection in real-world open scenarios.
This work designs an anonymized speaker prosody-emotion encoder based on the denoising diffusion probabilistic model, which can recover emotional style from anonymized prosody, thereby improving the anonymization effect and naturalness of speech. Unlike previous methods that rely on external emotion encoders, this method directly utilizes the denoising diffusion probabilistic model to repair emotional states from anonymized prosody.
With the widespread application of audio and music generation models in content creation, speech interaction, and intelligent media, training data privacy leakage has gradually become a security risk worthy of attention. Among these, membership inference attacks aim to determine whether a candidate audio sample has participated in model training, serving as an important tool for evaluating generative model memorization behavior and data leakage risks. This presentation focuses on the membership inference problem in audio diffusion models, with particular emphasis on why traditional scoring methods struggle to stably distinguish training members from similar non-members under the interference of same-distribution non-member samples and low false-positive rate requirements. The presentation will introduce an analysis approach based on the internal response of the diffusion reverse process, characterizing the model's potential memory signals for training samples by observing the perturbation stability of candidate samples at intermediate denoising stages. On this basis, a conditional calibration mechanism for music structure matching is further introduced to construct local hard non-member reference distributions for each candidate sample, thereby reducing false judgments caused by style, rhythm, and structural similarities. Experimental results show that this approach can improve member identification capabilities under strict low false-positive rate scenarios, providing new technical references for privacy evaluation of audio generation models, training data auditing, and responsible model release.
Xuechen Liu (Xi'an Jiaotong-Liverpool University): Xuechen.Liu@xjtlu.edu.cn
Shengchen Li (Xi'an Jiaotong-Liverpool University): Shengchen.Li@xjtlu.edu.cn
Meineng Zhu (University of International Relations): zmneng@uir.edu.cn