Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training. The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text. Existing methods then collapse this pair into a single scalar score with a fixed rule (geometric mean, weighted average) before taking the argmax. Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN). Each modality's standard representation becomes the real part of our pipeline, and a paired companion stream supplies the imaginary part. We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams. This results in two complex similarities, which are then fused. While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained. The four-stream complex architecture sets a new state of the art on both OV-AVEL benchmarks. On the open (unseen-class) split of OV-AVEBench we reach 66.5/59.1/54.1% Acc/Seg-F1/Event-F1 (+1.6/+4.1/+6.6 over the previously reported fine-tuned baseline), with consistent gains for seen classes as well. We also modify AVE dataset for this task and observe that our architecture reaches 60.7/51.9/50.4% Acc/Seg-F1/Event-F1, achieving state-of-the-art OV-AVEL results on it as well. We also propose a two-stream alternative, which also sees great improvements over the baseline.
Figures & tables
Figure 1: (a) 2- and 4-stream complex-fusion pipelines. (b) Unseen-split Segment-F1 vs Event-F1 on OV-AVEBench; the 4-stream complex model improves the baseline by +4.1 / +6.6 .
Figure 2: Four-stream complex-valued OV-AVEL architecture. Each modality forms a complex similarity tensor against the text prompts: Zv=Svt+iSvt\textsci (RGB +i⋅ iHSV, via frozen Ev ) and Za=Sat+iSpt (mel +i⋅ CycleGAN-translated phase, via frozen Ea and frozen GIF→mel ). The two tensors are stacked as the two channels of a five-layer ComplexCNN whose output magnitude gives the per-segment class scores. Only the four temporal-attention heads ( Φv,Φv\textsci,Φa,Φϕ ) and the ComplexCNN are trained; the ImageBind encoders and the CycleGAN generator are frozen
Figure 3: Unpaired CycleGAN translating the instantaneous-frequency (IF) phase spectrogram into the mel-magnitude domain expected by the frozen audio encoder. (a) forward and (b) reverse cycles, built from the two ResNet-9 generators GIF→mel , Gmel→IF and PatchGAN discriminators Dmel,DIF , trained with adversarial, cycle-consistency, identity, and distribution-matching ( Ldist ) losses. At OV-AVEL time only GIF→mel is used; the rest is training-time scaffolding.
Figure 4: The four input streams forming our two complex channels , shown for four event classes.
Table 1: Results on OV-AVEBench (%). The \textcolor red best and \textcolor blue second best results are highlighted in red and blue, respectively. The finetune-baseline and the previous works are from [ 32 ] .
Seen
Unseen
Total
Method
Acc
Seg
Eve
Avg
Acc
Seg
Eve
Avg
Acc
Seg
Eve
Avg
Training-free [ 32 ]
43.3
32.7
19.9
32.0
40.1
30.4
18.8
29.8
42.6
32.2
19.7
31.5
Fine-tuned baseline [ 32 ]
82.3
74.6
70.5
75.8
58.0
46.7
40.8
48.5
77.2
68.4
63.9
69.8
Ours (2-stream complex)
87.8
82.7
82.8
84.4
57.0
49.7
48.1
51.6
81.2
75.7
75.3
77.4
Ours (4-stream + ComplexCNN)
88.0
82.9
83.1
84.7
60.7
51.9
50.4
54.4
82.2
76.4
76.1
78.2
Table 2: Results on AVE-OV. All values in %. The same architecture wins on every split. On the unseen split, the 4-stream complex fusion improves the re-implemented fine-tuned baseline by +2.7/+5.2/+9.6/+5.9 points (Acc/Seg/Eve/Avg), and improves the 2-stream complex variant by +3.7/+2.2/+2.3/+2.8 .
Figure 6: Per-segment predictions on unseen-class videos. Two examples side-by-side; each panel shows the 10 video frames, the time-aligned audio envelope, and four rows: ground-truth labels, the OV-AVEL fine-tuned baseline [ 32 ] (geometric-mean fusion, trained temporal-attention heads), our 2-stream complex fusion, and our 4-stream complex fusion.
Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint audio-visual representations in Euclidean space, but still face two significant challenges. First, the lack of supervision signals for unseen categories makes it difficult to maintain audio-visual consistency across multiple temporal scales. Second, the lack of hierarchical constraints between segment- and video-level semantics prevents the model from establishing semantic consistency across different levels. To address these challenges, we propose a hierarchical semantic constrained heterogeneous graph (HSCHG) for audio-visual event localization framework. We first construct a heterogeneous hierarchical graph in Euclidean space, which includes audio and visual segment nodes and their corresponding video-level nodes. We use multi-directional temporal edges to capture complete temporal information within each modality. Simultaneously, we employ a dual-threshold filtering gated fusion strategy, introducing cross-modal information only when the alignment confidence is high. Furthermore, we introduce bidirectional semantic constraints between segment- and video-level representations to achieve semantic consistency across different levels. Based on this, we map the multi-level audio-visual representations and text prototypes uniformly into hyperbolic space. We use a hierarchical entailment regularization loss to characterize the hierarchical relationships between videos and segments. Extensive experimental results show that our method outperforms existing methods on the OV-AVEL benchmark. Ablation studies further validate the effectiveness of our method.
Zhe Yang, Ruyi Zhang, Hongtao Chen +4
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China · Peng Cheng Laboratory, Shenzhen 518000, China · Harbin Institute of Technology Suzhou Research Institute, Suzhou 215000, China
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At the intersection of these paradigms lies the task of Open-Vocabulary Audio Event Grounding: predicting all time intervals of a target sound event described by an arbitrary natural language query. While this task is crucial for real-world audio understanding and LALM adaptation, it is bottlenecked by data scarcity. Few large-scale resources provide open-vocabulary onset/offset supervision, and manual temporal annotation is prohibitively expensive. To address this, we introduce Auto-AEG, a scalable pipeline that constructs such supervision by automatic data construction and model fine-tuning. It pairs programmatically synthesized clips, which carry exact ground-truth intervals for supervised cold-start, with multi-model pseudo-labels on real-world audio that supply the reward signal for reinforcement learning. Training with this pipeline yields promising performance gains on both the DESED SED benchmark and AEGBench, an independent difficulty-stratified benchmark we release. Our results show that automatically constructed data, coupled with interval-aware reward function design, is an effective data-side route to expanding the temporal localization capability of LALMs.