Audio-Visual Semantic Segmentation (AVSS) aims to identify, segment, and classify sound-emitting objects in video frames. Previous Transformer-based AVSS approaches largely inherit design principles from image segmentation models. Recent studies show that these image segmentation models contain redundant components that contribute little to the segmentation performance. Following this insight, we propose Encoder-only Audio-Visual Segmentation (EASE). EASE runs at up to 365 FPS, 3x faster than prior State-of-the-Art (SotA) AVS models at comparable accuracy, and trains in under 11 GPU-hours. Furthermore, we achieve SotA AVSS performance across different backbones and input resolutions. Our results demonstrate that AVSS can be both simpler and faster, providing a scalable foundation for future research and real-time applications. Code, model weights, and samples are available at https://ease-avs.notion.site
Figures & tables
Figure 1 : Comparison between our simplified approach and common Transformer-based AVS methods. Our approach eliminates unnecessary architectural complexity, resulting in a more efficient model while preserving high segmentation performance.
Figure 2 : Overview of EASE. Audio and visual frames are encoded, and the encoded audio is enhanced by soft clustering it to semantically representative auditory centers. The query generator is used to generate sparse audio queries before concatenating them with the visual features. The concatenated sequence is processed with the remaining L2 encoder blocks and fused at each block with the enhanced audio features. The connection between enhanced audio features and AV Fusion is omitted for clarity. Finally, a mask module is employed to produce the final mask predictions from the visual tokens and the class predictions from the sparse audio query tokens.
AVSS
Method
Backb.
Params
Inp.
J↑
F↑
FPS ↑
GFLOPs ↓
AVSegFormer [ 13 ]
PVTv2
186M
2242
36.7
42.0
27
99
AVSegFormer [ 13 ]
PVTv2
186M
5122
37.3
42.8
23
504
AAVS [ 26 ]
Swin-B
187M
3842
48.5
53.2
69
151
Selm [ 20 ]
Swin-B
186M
4482
41.3
46.9
71
258
COMBO [ 37 ]
PVTv2
499M
2242
42.1
46.1
73
346
Table 1 : Quantitative comparison on AVSS [ 39 , 38 ] with efficiency evaluations. We compare performance on AVSS and throughput between EASE and SotA approaches with open-source code. † Retrained on the standard data split and benchmark metrics. ‡ DDESeg uses HTSAT (29M) as its audio encoder, while all other methods use VGGish (72M) [ 15 ] . Parameter counts are reported inclusive of audio encoders.
AVSS
S4
MS3
Method
Backb.
Inp.
J↑
F↑
J↑
F↑
J↑
F↑
TPAVI [ 39 ]
PVTv2
2242
29.8
35.2
78.7
87.9
54.0
64.5
CATR [ 21 ]
PVTv2
2242
32.8
38.5
84.4
91.3
62.7
74.5
ECMVAE [ 28 ]
PVTv2
2242
-
-
81.7
90.1
57.8
70.8
AQFormer* [ 17 ]
PVTv2
2242
-
-
81.6
89.4
62.2
72.7
AVSegFormer [ 13 ]
PVTv2
2242
36.7
42.0
82.1
89.9
58.4
69.3
Table 2 : Quantitative comparison on S4 , MS3 and AVSS categorized by backbone model and input size [ 39 , 38 ] . * MS3 model was pretrained on S4 data. † Retrained using standard data split and benchmark metrics. ‡ DDESeg uses HTSAT as its audio encoder, while all other methods use VGGish [ 15 ] .
Table 3 : Effect of pretraining on MS3. We pretrain EASE with PVTv2 [ 35 ] on S4 data and fine-tune it using MS3.
AVSS
Audio Enhancer Type
J↑
F↑
Linear
44.3
49.3
Precomputed Clusters
43.9
48.8
Learned Clusters
45.8
50.7
Table 4 : Ablation on different audio enhancers. EASE performance using DINOv3-B [ 33 ] 2242 model on AVSBench-Semantic [ 38 ] benchmark.
AVSS
Fusion Module Type
J↑
F↑
Bi-way Attention
43.5
48.7
AVFusion w/o Clustering
44.4
49.5
AVFusion
45.8
50.7
Table 5 : Ablation on different Audio-Visual Fusion Modules. EASE performance using DINOv3-B [ 33 ] 2242 model on AVSBench-Semantic [ 38 ] benchmarks.
Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic computational cost, limiting their suitability for resource efficient deployment. Most efficiency oriented methods focus on backbone reduction and overlook the interaction module as the primary bottleneck. This paper proposes LightAVSeg, a lightweight framework that replaces heavy attention with a decoupled design for semantic filtering and spatial grounding, resulting in interaction costs that scale linearly with spatial resolution. Furthermore, we introduce an auxiliary alignment loss to enforce semantic consistency during training with zero inference overhead. Extensive experiments demonstrate that LightAVSeg achieves a new state-of-the-art among lightweight methods: with 20.5M parameters ~1/7 of AVSegFormer), it reaches 50.4 mIoU on the MS3 benchmark and enables efficient inference on a mobile processor.
Qing Zhong, Guodong Ding, Lingqiao Liu +3
College of Informatics, Huazhong Agricultural University, Wuhan, China · School of Computing, National University of Singapore, Singapore · School of Computer Science, Adelaide University, Australia +2
Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio-visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic-Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio-modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8%. The code will be open-sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.
Leiye Liu, Miao Zhang, Jiahong Jiang +7
Dalian University of Technology Dalian, Liaoning, China · Carnegie Mellon University Pittsburgh, Pennsylvania, USA · Yale University New Haven, Connecticut, USA
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.