This special session aims to break down the barriers between independent research on speech and music, focusing on their deep coupling and dynamic interaction in multimodal human-computer communication. This session brings together scholars from auditory cognitive neuroscience and interdisciplinary fields to jointly explore scientific issues such as shared acoustic representations, rhythm perception and neural oscillation synchronization, cross-domain emotional modulation, and the constraints of consciousness on feature integration. Through interdisciplinary exchange, we expect to reveal the general computational principles of the auditory system in processing complex acoustic signals, provide neurophysiological evidence for cognition-inspired audio foundation models, deepen the understanding of the basis of human auditory cognition, provide theoretical support for immersive emotional empathic human-computer communication, and promote the leap of interactive experience from "semantic decoding" to "cognitive resonance."
Speech and music vary across cultures, but recent studies have found many cross-linguistic consistencies between speech and music, including the basic auditory rhythms of speech and music. This presentation will discuss whether the rhythms of speech and music reflect the inherent rhythms of the brain. Based on a review of previous research, two studies will be highlighted. One study involves comparing the relationship between the rhythms of speech and music and the rhythms of instinctive sounds such as crying and laughing; the other study explores whether the rhythms of speech and music can spontaneously emerge in simple sensorimotor synchronization experiments.
Conscious awareness requires establishing coherent perceptual representations of basic sensory features. Yet, whether consciousness is necessary for initiating the integration of basic sensory features remains unclear. Competing theories implicate distinct functional regimes of consciousness in the process of feature binding and creating conscious percepts. We used a novel multi-feature oddball paradigm with intracranial stereo-electroencephalography (sEEG) recordings in awake and anesthetized states to investigate the functional boundary of conscious awareness. In the awake state, the auditory attributes of loudness and tone, as well as the binding of the two features, were automatically encoded without attention to the stimulus. Moreover, anesthetically influenced cortical processes after stimulus offset. The results reveal the borderline of experience of conscious awareness constrains the feedforward and recurrent process directly at local rather than global level computations.
Yi Du (Institute of Psychology, Chinese Academy of Sciences): duyi@psych.ac.cn
Nai Ding (Zhejiang University): ding_nai@zju.edu.cn