Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph
Authors: Ziyang Chen, Yongjun Zhang, Wenting Li, Bingshu Wang, Yong Zhao, C. L. Philip Chen
Organizations: College of Computer Science, the State Key Laboratory of Public Big Data, Guizhou University, Guiyang 550025, China · School of Information Engineering, Guizhou University of Commerce, Guiyang 550021, China · School of Software, Northwestern Polytechnical University, Xi’an, 710129, China · Key Laboratory of Integrated Microsystems, Shenzhen Graduate School, Peking University, Shenzhen 518055, China · School of Computer Science and Engineering, South China University of Technology, Guangzhou 510641, China
Real-world applications of stereo matching, such as autonomous driving, place stringent demands on both safety and accuracy. However, learning-based stereo matching methods inherently suffer from the loss of geometric structures in certain feature channels, creating a bottleneck in achieving precise detail matching. Additionally, these methods lack interpretability due to the black-box nature of deep learning. In this paper, we propose MoCha-V2, a novel learning-based paradigm for stereo matching. MoCha-V2 introduces the Motif Correlation Graph (MCG) to capture recurring textures, which are referred to as ``motifs" within feature channels. These motifs reconstruct geometric structures and are learned in a more interpretable way. Subsequently, we integrate features from multiple frequency domains through wavelet inverse transformation. The resulting motif features are utilized to restore geometric structures in the stereo matching process. Experimental results demonstrate the effectiveness of MoCha-V2. MoCha-V2 achieved 1st place on the Middlebury benchmark at the time of its release. Code is available at https://github.com/ZYangChen/MoCha-Stereo.
Figures & tables
Fig. 1: In the deep learning process, feature channels are extended to higher dimensions. During this process, the geometric details learned in individual feature channels are inevitably lost. We propose that these details can be restored by identifying recurrent geometric structures across the channels. MoCha-V2 constructs a directed graph, i.e., motif correlation graph (MCG), to capture frequent patterns in feature channels. Specifically, MCG leverages node weights to detect recurring geometric structures within the channels. MCG establishes a robust and stable framework for identifying repetitive geometric structures.
Fig. 2: Architecture overview of MoCha-V2. First, a feature network extracts multi-scale features from the input stereo pair, and a channel-wise Haar wavelet transform decomposes the features into four subbands. A sliding window partitions each subband into non-overlapping 3×3 patches, which are treated as candidate motif subsequences. Second, for each channel group and window location, we construct a Motif Correlation Graph (MCG) in which nodes are these subsequences and directed edges are determined by Euclidean nearest-neighbour relationships; each node casts one unit vote to its nearest neighbours, and the vote-weighted mean of subsequences gives the motif maps. After applying the inverse wavelet transform and a sigmoid activation, each group yields a single motif channel. Third, the motif channels are broadcast and element-wise multiplied with the original vanilla channels, so that non-recurring patterns are attenuated while recurring geometric details are retained; this restoration step is parameter-free and thus fully interpretable. Fourth, the restored features are used to build a motif-weighted group-wise correlation volume, from which a differentiable soft argmin regresses an initial disparity map. The initial disparity is then refined by an iterative ConvLSTM update operator and further corrected by the Reconstruction Error Motif Penalty (REMP) module.
Table
Test benchmark (split)
Training set
III , VII - IX
SceneFlow (test)
S1: SF
III
Middlebury V3
S1: SF+TA+CR+FT+IS2+CH+MB
(test)
S2: CR+FT+IS2+CH+MB
IV
KITTI 2012
S1: SF
(reflective test)
S2: K15+K12
V
KITTI 2012
S1: SF
TABLE I: Training and testing protocols used for each result table. For every benchmark only its official training split is used for fine-tuning, and the test split used for evaluation is never seen during training.
Method
EPE (px)
D1>1px (%)
PSMNet [ 46 ]
1.09
12.1
RAFT-Stereo [ 35 ]
0.60
6.5
DLNR [ 25 ]
0.48
5.4
IGEV-Stereo [ 18 ]
0.47
5.3
MC-Stereo [ 47 ]
0.45
5.0
Selective-IGEV [ 27 ]
0.44
5.0
TABLE II: Quantitative evaluation on Scene Flow test set. The best result is bolded , and the second-best result is underlined . “EPE” means end-point error. “D1>1px” means percentage of stereo disparity outliers in first frame, and the error threshold is 1px.
Fig. 3: Visual comparisons with state-of-the-art stereo methods [ 53 , 27 ] on the Middlebury test set. We present the disparity map and error map with a 1-pixel threshold. In the error map, the black regions indicate areas of misestimation.
Method
D1>2px (%) ↓
D1>3px (%) ↓
D1>4px (%) ↓
Noc
All
Noc
All
Noc
All
DPCTF-S [ 56 ]
9.92
12.30
6.12
7.81
4.47
5.70
CREStereo [ 57 ]
9.71
11.26
6.27
7.27
4.93
5.55
NLCA-Net V2 [ 58 ]
11.20
13.19
6.17
7.65
4.06
5.19
UCFNet [ 52 ]
9.78
11.67
5.83
7.12
4.15
4.99
IGEV-Stereo [ 18 ]
7.29
8.48
4.11
4.76
2.92
3.35
TABLE IV: Results on the KITTI 2012 Reflective leaderboard. Bold : best performance, underline : second best.
Fig. 4: Visual comparisons with state-of-the-art stereo methods [ 52 , 27 ] on the KITTI 2012 test set. In the first row of images, the streetlight presents a matching challenge due to its thin structure. UCFNet fails to capture the presence of the poles, and Selective-IGEV inaccurately estimates the sky above the streetlight as being closer in depth to the streetlight. Only MoCha-V2 provides a more accurate estimation of the streetlight’s edges. In the second row, the roofs of the two cars in the image reflect natural light, which affects the accuracy of stereo matching. Both UCFNet and Selective-IGEV are impacted by this reflection and fail to accurately capture the cars’ edge textures. Additionally, Selective-IGEV does not estimate the contours of the poles on the wall. In the third row, due to lighting influences, UCFNet and Selective-IGEV misinterpret the buildings in the background as part of the trees in the foreground. In contrast, MoCha-V2 delineates a more accurate edge texture.
Method
KITTI 2015 (%) ↓
KITTI 2012 (%) ↓
D1-all
D1-fg
D1-bg
All
Noc
CREStereo [ 57 ]
1.69
2.86
1.45
2.18
1.72
NLCA-Net V2 [ 58 ]
1.77
3.56
1.41
2.34
1.83
CroCo-Stereo [ 51 ]
1.59
2.65
1.38
-
-
DLNR [ 25 ]
1.76
2.59
1.60
-
-
IGEV-Stereo [ 18 ]
1.59
2.67
1.38
2.17
1.71
TABLE V: Results on the KITTI 2015 and KITTI 2012 leaderboards. In the KITTI 2015 table, “D1” means percentage of stereo disparity outliers in first frame, “all” means percentage of outliers averaged over all ground truth pixels, “fg” means percentage of outliers averaged only over foreground regions, “bg” means percentage of outliers averaged only over background regions. In the KITTI 2015 table, “All” in KITTI 2012 means ercentage of erroneous pixels in total, “Noc” means percentage of erroneous pixels in non-occluded areas. Error threshold is 2 px for KITTI 2012.
Fig. 5: Visual evaluations using the KITTI 2015 test set in contrast to the SOTA techniques [ 52 , 27 ] . In the first row, UCFNet and Selective-IGEV failed to detect the ropes along the roadside. In the second row, Only our method successfully detects the presence of the signage.
Method
Noc (%) ↓
All (%) ↓
RAFT-Stereo [ 35 ]
7.04
7.33
CREStereo [ 57 ]
3.58
3.75
IGEV-Stereo [ 18 ]
3.52
3.97
GMStereo [ 53 ]
5.94
6.44
CREStereo++ [ 50 ]
4.61
4.83
LoS [ 23 ]
3.59
3.83
TABLE VI: Quantitative Comparisons with SOTA Methods on the ETH3D Benchmark. Error threshold is 0.5 px.
Fig. 6: Visual comparisons with SOTA stereo methods [ 53 , 60 ] on the ETH3D test set. In the first row, GMStereo and GANet+ADL fail to capture the fine-level geometry of sluices and pipes. In the second row, reflection effects make it challenging for existing methods [ 53 , 60 ] to accurately identify the contours of the sculpture. Furthermore, among the three methods, only MoCha-V2 successfully detects the outline of the thin object adjacent to the sculpture. In the third row, lighting effects mislead existing methods, resulting in inaccuracies when identifying lamps, walls, and thin objects. In contrast, MoCha-V2 accurately identifies these objects and preserves their geometric contours.
No.
Model
REMP
MCA
MCG
EPE (px)
D1>1px (%)
Time (s)
1
Baseline
0.451
5.247
0.34
2
MoCha-Stereo [ 20 ]
✓
✓
0.409
4.851
0.35
3
MoCha-Stereo w/ WT w/o G
✓
◯
0.401
4.825
0.35
4
MoCha-V2
✓
✓
0.381
4.304
0.32
TABLE VII: Ablation study for MoCha-V2. The baseline employed in these experiments utilized EfficientNet [ 37 ] as the backbone for IGEV-Stereo [ 18 ] with 32 iterations. The Time denotes the inference time on single NVIDIA A100. “WT” means Wavelet Transform in our MoCha-V2, “G” means Gaussian filter in our MoCha-Stereo. The Time denotes the inference time on single NVIDIA A100.
Fig. 7: Visualization of Motif Correlation Graphs computed from a pair of images in the Scene Flow dataset. The darker the color of a node, the greater its weight. Arrows point from each node to the node with the smallest Euclidean distance to it.
Fig. 8: Visualization of feature channels. We utilize Principal Component Analysis (PCA) to reduce the channel features to a one-dimensional feature map and visualize it.
Method
Number of Iterations
1
2
4
8
16
32
IGEV-Stereo [ 18 ]
0.66
0.62
0.55
0.50
0.47
0.47
Selective-IGEV [ 27 ]
0.65
0.60
0.53
0.48
0.45
0.44
MoCha-Stereo [ 20 ]
0.56
0.52
0.46
0.42
0.41
0.41
MoCha-V2 (Ours)
0.49
0.45
0.42
0.40
0.39
0.38
TABLE VIII: Ablation study for iterations. The metric here is EPE (px). Bold : Best Performance.
EPE (px)
Time (s)
Iters
[ 20 ]
V2
±
[ 20 ]
V2
1
0.56
0.49
+ <1%
0.19
0.17
2
0.52
0.45
+ 1.9%
0.20
0.18
4
0.46
0.42
+ 2.2%
0.22
0.20
8
0.42
0.40
+ <1%
0.26
0.23
16
0.41
0.39
+ 2.4%
0.30
0.27
TABLE IX: A comparison of accuracy and speed between MoCha-Stereo and MoCha-V2. The “+” symbol indicates the percentage by which MoCha-V2 outperforms MoCha-Stereo, while the “-” shows the reverse. “Iters” means the number of iterations. “Time” refers to the inference time on a single NVIDIA Tesla A100.
School of Control Science and Engineering, Dalian University of Technology, Dalian, 116024, China · Pika Labs · Department of Automation, Shanghai Jiao Tong University, Shanghai, 200240, China +1