Spatial hearing and human-robot interaction for embodied intelligence constitute a critical foundation for robots to enter real-world environments and achieve autonomous human-robot collaboration. Analogous to how humans and other vertebrates rely on auditory systems to perceive their surroundings, embodied intelligent agents must begin with auditory perception to recognize, localize, and track sound events, subsequently performing auditory scene analysis, spatial mapping, and target source selection, thereby driving navigation, collaboration, and human-robot interaction tasks. Moreover, auditory and spoken language interaction serves as a vital modality in human social communication. Real-world human-machine communication must extend beyond mere speech clarity and naturalness to encompass the comprehension of multi-speaker contexts, spatial relationships, interaction protocols, implicit emotions, and social norms.
This session invites discussion on the auditory perception and human-robot interaction of intelligent agents in commercial exhibition, home companionship, and industrial collaboration scenarios, including but not limited to:
The ability to localise, track, and selectively attend to sound sources in dynamic environments relies on the seamless integration of auditory and visual information. This talk reviews our work on audio-visual signal processing for egocentric perception, spanning sound source localisation, tracking, and selective attention in dynamic environments. The focus will be on the seamless integration of auditory and visual information, with applications to wearable spatial intelligence, robotics, and augmented reality.
Replicating this capability in machines—enabling them to localise, track, and selectively attend to sound sources in dynamic environments—has broad applications in wearable spatial intelligence, robotics, and augmented reality. The talk reviews our work on audio-visual signal processing for egocentric perception, including sound source localisation, tracking, and selective attention in dynamic environments. Recent developments in multimodal large language models and their integration with audio-visual perception will also be discussed, with future directions towards unified egocentric spatial intelligence.
This report aims to explore "Embodied Audition" as an independent core perceptual capability of embodied intelligence. Traditional auditory research is often confined to the passive framework of signal processing, where the intelligent agent acts as a bystander to environmental sound sources. Under the embodied intelligence paradigm, we propose that audition is dynamically generated "behavioral feedback" during agent interaction. By actively closing the loop between audition and agent actions, sound is no longer merely information data, but a key for the agent to parse physical properties, spatial depth, and interaction dynamics. This report will argue why embodied audition should be regarded as a new track independent of traditional acoustics research, and outlook its irreplaceable role in building physical commonsense for intelligent agents.
Spatial audio provides users with immersive perception beyond traditional stereo by reconstructing three-dimensional auditory experiences, and has become a core enabler of immersive technologies such as virtual reality and augmented reality. This report first reviews the representation formats, understanding tasks, generation tasks, and related datasets and evaluation metrics of spatial audio in chronological order. Secondly, addressing the challenge of scarce high-quality multimodal spatial audio data, it introduces multimodal recording spatial audio datasets including audio spatialization, spatial speech synthesis, spatial singing voice synthesis, spatial music generation, and sound source localization and detection. On this basis, the report presents the OmniAudio task, which directly generates FOA (First-order Ambisonics) format spatial audio from 360° panoramic video. Finally, to address the insufficient real-time streaming generation capability, it introduces a causal autoregressive diffusion Transformer architecture and spatial video-audio contrastive learning strategy, achieving streaming synchronous generation of high-quality spatial audio from panoramic video and text prompts.