cs.CVAug 14, 2025

MAGIC: Learning from Visibility Asymmetry for Unsupervised Stereo Matching

Authors: Peng Xu, Zhiyu Xiang

Organizations: College of Information Science and Electronic Engineering (ISEE) Zhejiang University, China

Abstract

Learning disparity in occluded regions remains difficult for unsupervised stereo matching. Photometric supervision lacks valid target-view correspondences in these regions, while the teacher and student in conventional binocular self-training share the same target view and therefore the same occlusions. Even when supervision is available, the small proportion of occluded pixels limits their contribution to training. We propose MAGIC, a multi-baseline geometric consistency framework for reliable and effective occlusion supervision. The teacher and student share a reference image but use different target views, allowing the teacher to observe correspondences that are occluded from the student. After aligning disparities across baselines, MAGIC uses predictions from teacher-visible regions to supervise student-occluded regions. An occlusion-aware weighting strategy strengthens supervision on teacher-visible but student-occluded pixels, preventing their training signal from being overwhelmed by non-occluded regions. We also introduce MBS20K, a synthetic multi-baseline stereo dataset spanning diverse scenes, weather, and lighting. Pre-trained on MBS20K, MAGIC generalizes to real-world datasets with consistently fewer occluded-region outliers. On KITTI, the pre-trained model already outperforms several fine-tuned unsupervised methods. Fine-tuning this model on standard binocular pairs achieves state-of-the-art unsupervised performance on KITTI 2015 and 2012. Incorporating synthesized multi-baseline views during fine-tuning further improves performance. Our code and dataset will be released upon acceptance.

Explore similar work

Apr 22, 2026cs.CV

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot ability. However, due to the limitations of ViTs in handling resolution sensitivity and their relative neglect of local information, the ability of ViT-based methods to predict details and handle arbitrary-resolution images is still weaker than that of CNN-based methods. To address these shortcomings, we propose MLG-Stereo, a systematic pipeline-level design that extends global modeling beyond the encoder stage. First, we propose a Multi-Granularity Feature Network to effectively balance global context and local geometric information, enabling comprehensive feature extraction from images of arbitrary resolution and bridging the gap between training and inference scales. Then, a Local-Global Cost Volume is constructed to capture both locally-correlated and global-aware matching information. Finally, a Local-Global Guided Recurrent Unit is introduced to iteratively optimize the disparity locally under the guidance of global information. Extensive experiments are conducted on multiple benchmark datasets, demonstrating that our MLG-Stereo exhibits highly competitive performance on the Middlebury and KITTI-2015 benchmarks compared to contemporaneous leading methods, and achieves outstanding results in the KITTI-2012 dataset.
Jul 6, 2026cs.CV

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bottleneck. Training-based methods achieve superior quality but rely on scarce stereo pairs or synthetic data with domain gaps. We address this through the first self-supervised framework learning from monocular videos via cycle consistency. Our key contribution is the Geometric Reciprocity Theorem (GRT): under the nearest-neighbor DIBR formulation, the disocclusion mask when synthesizing a target view equals the mask of pixels lost when warping back from target to source, enabling analytical computation of test-time disocclusion masks directly from monocular images. This yields train-test consistency for the stated warping formulation, supporting self-supervised learning from unlimited monocular videos and substantial improvements over training-free and supervised state-of-the-art methods. Project page: https://visual-ai.github.io/grt/
Jul 5, 2026cs.CV

AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation

Learning-based stereo matching models struggle in underwater environments due to scarce in-domain data and the difficulty of extracting discriminative correspondences from degraded imagery. In this work, we present AquaStereo\textbf{AquaStereo}, a perception-enhanced framework with a data simulation pipeline and a self-distillation strategy that jointly address data scarcity and feature degradation in underwater stereo matching. First, a depth-conditioned diffusion pipeline renders underwater stereo pairs while preserving binocular geometry, with a lightweight left-right consistency module ensuring geometric alignment. Training on this synthetic corpus effectively narrows the terrestrial-underwater gap and improves zero-shot robustness. Second, a frozen binocular teacher trained on clean terrestrial pairs guides a student exposed to rendered underwater pairs with perturbations. A stage-weighted sequence loss is performed to align the student's disparities with the teacher's geometry, while a clean-branch supervision with shared pseudo targets prevents scale drift. To further enhance feature stability under turbidity and low texture, we introduce learnable perception frames, a perception-enhanced feature formulation that constructs robust matching descriptors by fusing temporal cues from two auxiliary views encoded by a video backbone with semantic features extracted by a strong image encoder. Extensive experiments demonstrate that AquaStereo\textbf{AquaStereo} substantially improves robustness and zero-shot generalization in challenging underwater scenarios. The code is available at https://github.com/qz-wei/AquaStereo.