Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL-based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly-supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully-supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at https://github.com/fraunhoferhhi/RIPEpp .
Figures & tables
Figure 1 : We propose a reward formulation for weakly-supervised keypoint extraction training that enables training from positive pairs only and is directly applicable to the training of learned matching architectures, demonstrated with LightGlue [ 25 ] .
Figure 2 : Schematic overview of the proposed weakly-supervised keypoint extraction framework. During training, the framework operates on pairs of images, from which a neural network generates heatmaps H from which keypoint locations K with a probability p are sampled. Descriptors D are extracted from the encoder, as Hypercolumn Features. Using these descriptors, keypoints are matched and filtered using epipolar constraints to calculate the reward matrix R , which weights the combined log-probabilities L of the keypoint locations, resulting in an approximation of the gradients, using REINFORCE [ 49 ] .
Figure 3 : Visual comparison of the reward matrix construction to shape the log-probabilities in L for the gradient computation ( Eq. 1 ). In RIPE (left) rewards are assigned at image pair level: geometric consistent matches are rewarded for positive pairs and penalized for negative sample pairs, outliers are ignored. In RIPE ++ (right), rewards are defined at correspondence level within positive pairs only. Geometrically consistent matches are rewarded and geometrically inconsistent matches are explicitly penalized. This finer-grained reward makes negative pairs unnecessary and yields a more informative training signal.
Figure 4 : Schematic overview of the proposed weakly-supervised training framework for LightGlue. Given keypoint locations and descriptors from two images, LightGlue refines the representations through multiple layers of self- and cross-attention, incorporating positional and cross-image context. Rather than supervising with ground-truth correspondences derived from depth or pose data, we formulate matching as a reinforcement learning problem: matches deemed geometrically consistent by RANSAC receive a positive reward, while outliers are penalized. This enables end-to-end training of the matcher using only image pairs.
Method
AUC @5∘
AUC @10∘
AUC @20∘
Supervision
Geometric GT
ALIKED [ 53 ] TIM’23
56.66
69.64
79.35
Pose+Homog.
DaD [ 12 ] CoRR’25
56.46
70.07
80.18
Depth
DeDoDe-B [ 13 ] CVPRW’24
56.01
69.20
78.72
Pose
DISK [ 45 ] NeurIPS’20
50.69
64.35
74.74
Pose/ Depth
RACO [ 43 ] 3DV’26
57.96
71.28
81.00
Homog. (Det.) + Pose (Desc.)
No geo. GT
SuperPoint [ 8 ] CVPRW’18
47.26
60.89
70.65
Homography
Table 1 : Relative pose estimation on MegaDepth1500. Methods are grouped by whether training requires geometric ground truth (pose, depth) for the detector or descriptor. RIPE ++ trains both detector and descriptor from positive image pairs alone, without geometric supervision, yet remains comparable to geometric ground truth group and outperforms RIPE despite discarding its negative pairs. Best and second-best are highlighted within each group.
Figure 5 : Qualitative results of raw matches for RIPE ++ and RIPE ++ in combination with our weakly-supervised LightGlue matcher on the MegaDepth1500 benchmark dataset on the left and right respectively. The middle image shows qualitative results for RIPE ++ Medical on the SCARED1500 test dataset. For more qualitative results please refer to the supplementary material.
Method
AUC @5∘
AUC @10∘
AUC @20∘
Supervision
ALIKED [ 53 ] TIM’23
18.24
40.61
61.23
Pose+Homography
DaD [ 12 ] CoRR’25
18.44
40.56
61.40
Depth
DeDoDe-B [ 13 ] CVPRW’24
18.85
41.69
62.80
Pose
DISK [ 45 ] NeurIPS’20
16.58
36.77
55.63
Pose/ Depth
RACO [ 43 ] 3DV’26
19.35
42.36
63.34
Homography
SIFT [ 27 ] IJCV’04
13.24
32.17
51.66
—
Table 2 : Evaluation of the relative pose estimation on Scared 1500, showing the benefits of our proposed training scheme. It allows to easily train a specialized keypoint extractor, simply from video frames. The best and second-best performances are highlighted.
Parameter
Metrics
pos only
entropy
# RANSAC inlier
% RANSAC inlier
AUC @5∘
AUC @10∘
AUC @20∘
–
297.5
34.4
51.83
65.37
75.94
✓
–
427.5
41.5
52.42
66.78
78.03
✓
1\times10−4
–
–
–
–
–
✓
1\times10−5
266
33.6
49.62
62.7
72.8
✓
1\times10−6
366.5
40.6
56.59
69.44
79.18
Table 3 : Ablation study of our proposed methods on the MegaDepth1500 test set.
Figure 6 : Effect of our entropy-based regularization on keypoint localization. Without it, heatmaps tend to become diffuse, particularly at low input resolutions (top-left). Regularization encourages sharper, more peaked responses (bottom).
Method
AUC @5∘
AUC @10∘
AUC @20∘
RIPE ++
56.58
69.53
79.33
↓ + LightGlue
+3.07
+4.16
+4.53
Matched-RIPE ++
59.65
73.69
83.86
Table 4 : Improvements on MegaDepth1500 from training LightGlue with our proposed approach on positive image pairs only.
Extractor
Matcher
Name
Value
Purpose
Name
Value
Purpose
ρin
1.0
reward Eq. 3
νin
1.0
reward matcher training
ρout
-0.1
penalty Eq. 3
νout
-1.0
penalty matcher training
λ
-1e-7
penalty Eq. 3
λ
-1e-07
penalty no match
ψ
5.0
weight Ldesc
η
0.0001
weight Lnm Eq. 10
ω
1e-6
weight LH
Table 5 : Overview of training hyperparameters for the keypoint extractor (left) and the matcher (right).
Figure 7 : Illustration Distance-based Reward
Parameter
Metrics
pos only
contrastive
infonce
distance reward
curriculum
entropy
# RANSAC inlier
% RANSAC inlier
AUC @5∘
AUC @10∘
AUC @20∘
✓
297
34
51.83
65.37
75.94
✓
✓
427
41
52.42
66.78
78.03
✓
✓
✓
384
40
54.75
68.33
78.92
✓
✓
✓
347
38
52.57
65.84
76.43
✓
✓
430
45
51.72
65.0
75.9
Table 6 : Ablation of the different components on the MegaDepth1500 benchmark set with the longer side of the input rescaled to 1200 pixels and 2048 keypoints detected.
Figure 8 : Number of false NN matches (left) and RANSAC inliers (right) for negative image pairs.
Method
Day
Night
.25m/ 2 °
.5m/ 5 °
5m/ 10 °
.25m/ 2 °
.5m/ 5 °
5m/ 10 °
ALIKED
85.5
92.0
95.8
66.0
80.6
92.7
DaD
88.9
93.2
97.0
69.6
85.9
96.9
DeDoDe
81.8
89.1
93.4
55.5
68.1
78.0
DISK
81.9
91.4
95.4
61.8
75.9
86.9
RaCo
87.4
93.4
97.1
72.8
86.9
96.9
Table 7 : Results on outdoor visual localization using Aachen Day-Night v1.1.
Figure 9 : Qualitative results on the MegaDepth1500 benchmark dataset for, RIPE [ 19 ] (top), RIPE ++ (ours, middle) and RaCo [ 43 ] (bottom).
Figure 10 : Qualitative results on MegaDepth1500 for RIPE ++ paired with our weakly-supervised LightGlue matcher.
Figure 11 : Qualitative results on the Scared1500 benchmark dataset for RIPE ++ (ours, top), RIPE ++ Medical (ours, middle) and RaCo [ 43 ] (bottom).
Chalmers University of Technology, Sweden · Mobile Perception Lab, ShanghaiTech University, China · LIGM, Ecole des Ponts, Univ. Gustave Eiffel, CNRS, France