MimicSound: Learning Bimanual Manipulation from Audio-Visual Human Videos

Abstract

Learning robot manipulation from human visual videos has shown promise, but leveraging audio information in human data remains underexplored. To address it, we propose MimicSound, a unified End2End audio-visual imitation learning framework that novelly co-trains on human audio-visual data together with robot demonstrations using a dual-branch design. Human-robot data collection employs a shared top-view camera and microphone to ensure spatial and auditory alignment, with human audio augmented by robot noise to reduce cross-domain gaps. To achieve temporal consistency, we introduce Timestep-aligned Audio Synchronization for per-step multimodal alignment in robot data, and Marker-based Human-Robot Demonstration Alignment for phase-wise speed alignment between the two data sources across long-horizon tasks. During training, MimicSound adopts a shared audio-visual encoder: the vision module encodes top-view robot images or hand-masked human images to mitigate domain bias, while the shared audio encoder, based on a pretrained Audio Spectrogram Transformer (AST) fine-tuned via Weight-Decomposed Low-Rank Adaptation (DoRA), extracts unified auditory representations for both branches. Two Bidirectional Cross-modal Attention Modules then dynamically fuse audio and visual features in each branch before feeding them into the Action Chunking with Transformers (ACT) policy. We evaluate MimicSound on three bimanual audio-visual tasks, demonstrating the advantage of learning from human audio-visual demonstrations, while ablations show the contribution of each component.


Approach Overview

Approach Overview

Fig. 1. Overview of the MimicSound. The framework adopts a dual-branch design for human and robot data. In the robot branch, multi-view RGB images $o_{R,t}^{cams}$ serve as visual inputs: the left and right arm views are encoded by ResNet models pre-trained on ImageNet, while the top-view image and audio $o_{R,t}^{aud}$ are processed by the shared audio-visual encoder. Robot joint states $o_{R,t}^{jnt}$ are normalized and encoded via a linear layer. The human branch takes the top-view RGB image $o_{H,t}^{cam}$, robot-noise-augmented audio $o_{H,t}^{aud}$, and normalized hand positions $o_{H,t}^{pos}$, also encoded by a linear layer. The shared encoder includes a vision and an audio module: top-view robot image or hand-masked human image is encoded by a shared ResNet, while audio is downsampled, converted into Mel spectrograms, and encoded by an AST initialized from AudioSet. During training, the AST backbone is frozen and only DoRA modules are fine-tuned. In each branch, audio and visual features are fused by their Bidirectional Cross-modal Attention module before being fed into the ACT-based policy. The policy predicts action chunks for robot joints and end-effector positions in the robot branch, and for hand positions in the human branch, where the latter shares the same linear layer as the robot end-effector prediction. Losses are computed separately for each branch (robot joints, robot end positions, and human end positions) and jointly optimized through a shared gradient update during training.


Human-Robot Audio-Visual Data Collection and
Autonomous Bimanual Skill Learning with Audio.

🔊 Please turn on your speaker for full experience