Tracking a novel object's 6D pose over long horizons currently requires either expensive onboarding or a reconstruction maintained throughout the sequence. This makes current trackers impractical for robotic manipulation and augmented reality, which need trackers that are ready to use and run in real time. We show that a lightweight tracking module can be applied on top of a wide range of correspondence estimation methods to keep drifts bounded while maintaining fast runtime. Our key idea is to avoid point-based optimization in the pose graph and operate only on relative pose constraints, which we weight by a derived uncertainty from the geometric alignment. This makes optimization independent of the number of correspondences while avoiding the direct inclusion of noisy point measurements, leading to fast and robust long-term tracking. Across four real-world benchmarks, our approach achieves tracking accuracy comparable to reconstruction-based trackers with a fraction of the optimization cost. Overall, these results suggest that a compact and reliable pose graph optimization can provide long-horizon consistency at substantially lower computational cost.
Figures & tables
Figure 1 : Left: We propose a lightweight generalizable 6D pose tracking framework for novel objects that can be flexibly paired with any correspondence estimation method. Right: We show that our approach achieves the best pose accuracy and runtime efficiency tradeoff on the YCBInEOAT dataset. Gray: reported in the original paper on different hardware, see Tab. 4 .
Figure 2 : High-level method overview. We initialize a SE(3) pose graph which will be incrementally updated when a new frame arrives. In the first stage of tracking ( 1 ), we select frames from the past which we want to connect to the current frame via edges in the pose graph. Next, for each pair of frames corresponding to an edge, we estimate their correspondences ( 2.a ), and use RANSAC Umeyama to solve 6D relative poses ( 2.b ) and estimate their uncertainty( 2.c ). Finally, we insert ( 3.a ) and optimize ( 3.b ) the new node and its edges to the pose graph using iSAM2 , and update the memory pool ( 3.c ).
Figure 3 : Detailed method overview. For a new frame Ii , we first select K frames from previous frames to build candidate loop closure constraints based on the geodesic distance and farthest frame sampling. Then the current frame Ii and selected K candidate frames will be passed to the correspondence estimation method and RANSAC Umeyama to obtain relative poses Tik and uncertainty Σik . Furthermore, we insert the new node Vi with its odometry and loop closure edges to the uncertainty-aware graph and update Ti using iSAM2. Last, we update the memory pool M based on viewpoint diversity.
Cat.
Method
HOI4D (300 frames)
YCBInEOAT (823 frames)
HO3D (1547 frames)
YCB Video (1718 frames)
Avg. Speed fps
ADD-S (%)
ADD (%)
ADD-S (%)
ADD (%)
AR (%)
ADD-S (%)
ADD (%)
AR (%)
ADD-S (%)
ADD (%)
AR (%)
CAD
FoundationPose (CAD)
84.4
66.5
96.4
93.1
90.5
97.0
88.9
89.3
98.0
96.0
92.8
22.9
RPE
UNOPose
92.2
79.4
91.0
67.7
65.7
92.9
51.6
31.7
96.4
84.7
80.6
10.0
One2Any
79.4
64.0
66.0
31.5
31.1
75.5
37.1
29.4
92.9
88.7
75.0
30.0
TbR
BundleSDF
90.8
79.0
93.8
87.0
82.2
96.5
92.6
87.9
97.6
93.8
90.0
1.8
6DOPE-GS †
-
-
93.8
87.8
-
95.1
84.3
-
-
-
-
4.2
Table 1 : Pose tracking accuracy and runtime across four benchmarks. We report ADD-S and ADD (AUC %, 0 – 0.1 m), and BOP Average Recall (AR) on HOI4D, YCBInEOAT, HO3D, and YCB-Video (average sequence length shown under each dataset name), alongside average inference speed (fps) across all four datasets. AR is omitted for HOI4D, because of no BOP-format ground truth available. Methods are grouped by category: CAD : requires CAD models; TbR : tracking-by-reconstruction; RPE : relative pose estimation; PGO : pose-graph-optimization-based tracking, the category our method belongs to alongside BundleTrack. Best result among non-CAD methods per column in bold , second-best underlined (tied values are marked identically). Among non-CAD methods, our method leads on YCBInEOAT, YCB-Video, and HOI4D, and is comparable to BundleSDF on HO3D, while running substantially faster than other tracking-by-reconstruction and pose-graph-based baselines. BundleTrack (R.): BundleTrack re-run with RoMa v2 as the feature matcher, using the official codebase. † : numbers taken from the original paper runtime reported on a different GPU than ours ( Tab. 4 )
Cat.
ADD (%)
ADD-S (%)
Ours
BS
BT
UP
FP
Ours
BS
BT
UP
FP
\faCarSide
86.3
84.4
84.9
83.2
60.0
93.5
92.9
93.3
93.2
72.8
83.3
83.3
83.5
80.9
78.3
92.5
92.5
92.6
92.2
89.7
82.4
78.2
81.2
82.5
75.1
91.7
87.5
90.9
92.2
84.8
\faChair
75.3
74.8
75.3
76.2
76.3
88.9
87.6
88.1
89.2
89.3
74.3
74.1
74.2
73.9
42.6
94.4
93.5
93.9
94.1
85.4
Table 2 : Per-category comparison on five HOI4D rigid object categories. We report ADD-AUC and ADD-S-AUC (%, 0 – 0.1 m). BS: BundleSDF, BT: BundleTrack, UP: UNOPose, FP: FoundationPose.
Matcher
Matching
Solve+Unc.
Frame Sel.
PGO
Pool Update
Ours (fps)
RoMa v2
0.05
0.04
5e-4
2e-3
2e-4
10.8
Table 3 : Per-stage latency on YCB Video. Results reported in seconds per frame and averaged over the dataset. The final column reports total tracking throughput in frames per second.
Method
GPU
fps
BundleTrack (LF-Net) [ 40 ]
RTX 2080 Ti
10.0
6DOPE-GS [ 18 ]
RTX 4090
4.2
Ours
RTX 4090
9.0
Table 4 : Runtime on different GPUs. Numbers are reported from the original papers.
Method
Matcher
HOI4D
YCBInEOAT
HO3D
YCB Video
ADD-S(%)
ADD(%)
ADD-S(%)
ADD(%)
ADD-S(%)
ADD(%)
ADD-S(%)
ADD(%)
BundleTrack
RoMa v2
91.8
80.2
93.8
88.0
96.5
92.6
97.5
94.9
Ours
92.2
80.3
94.4
88.8
96.4
92.2
98.2
96.5
BundleTrack
LoFTR
92.0
80.1
92.5
84.9
94.1
79.3
97.3
93.9
Ours
92.2
80.6
93.0
87.8
94.8
85.7
97.4
94.2
Table 5 : Comparison to dense alignment-PGO method BundleTrack. We show the direct comparison between dense-alignment and pose consistency optimization under the same correspondence estimator. The pose-level optimization preserves accuracy while dramatically reducing optimization cost.
Uncertainty
HOI4D
YCBInEOAT
HO3D
YCB Video
ADD-S(%)
ADD(%)
ADD-S(%)
ADD(%)
ADD-S(%)
ADD(%)
ADD-S(%)
ADD(%)
Identity ( RoMa v2 )
92.11
80.33
93.69
87.60
96.14
90.77
96.82
94.19
Trace
-0.28
-0.23
-1.15
-3.07
+0.39
+1.80
+1.34
+2.24
Covariance
+0.09
+0.00
+0.70
+1.20
+0.26
+1.43
+1.34
+2.31
Oracle
+0.40
+1.56
+1.08
+2.40
+0.80
+2.95
+0.17
+1.20
Identity ( LoFTR )
91.40
79.35
92.38
83.52
94.07
83.34
96.03
92.30
Table 6 : Ablation on uncertainty structures. We compare four levels of uncertainty information used in pose graph optimization: Identity (no per-edge uncertainty), Trace (isotropic, using only the scalar magnitude of our estimated covariance), Covariance (our full anisotropic per-edge estimate), and Oracle* (derived from ground-truth pose error, an upper bound).
Figure 4 : Drift analysis on YCB Video dataset. Rotation and translation drift over the full 1718-timestep sequence.
Figure 5 : Overconfident uncertainty for heavy occluded symmetric objects and thin objects from end-on views. Top: under the heavy occlusion, the hand covers the label, leaving only correspondences on the rotationally-symmetric body; Bottom: thin object viewed from its end-on view axis, leaving little visual signal. For the purpose of visualization, we show the square root of the covariance’s trace as a scalar uncertainty.
Figure 6 : Uncertainty vs. Pose Error. We show the joint distribution of estimated uncertainty and ground-truth pose error, for rotation (left) and translation (right), on HO3D and YCBInEOAT. Hexbin color indicates edge density on a log scale; the red line and shaded band show the median and interquartile range of error within log-spaced uncertainty bins.
Keyframe
HO3D
YCBInEOAT
ADD-S (%)
ADD (%)
ADD-S (%)
ADD (%)
Ours
96.37
92.23
94.18
88.75
w.o FFS
-4.22
-14.10
-1.71
-3.81
w.o Pool
-0.10
-0.29
-0.33
-0.56
Table 7 : Ablation on keyframe selection. Ours w.o. FFS use the direct K neighbors to build loop closure edges; Ours w.o. memory pool allows loop-closure candidates to be drawn from all past frames.
Figure 7 : Parameter Sweep of τk for different feature matchers.
Setting
PGO Target
RPE
HO3D
YCBInEOAT
ADD-S (%)
ADD (%)
ADD-S (%)
ADD (%)
F2F Tracking
-
UNOPose
20.1
5.8
22.1
10.2
LoFTR
63.4
30.5
36.7
24.1
RoMa v2
78.0
56.7
70.8
52.5
BundleTrack
Dense-Alignment
LFNet
92.4
66.0
93.0
87.3
LoFTR
94.1
79.3
92.5
84.9
Table 8 : Comparison of tracking approaches on HO3D and YCBInEOAT with different correspondence and pose estimation methods. Top : direct frame-to-frame (F2F) tracking using each correspondence/relative pose method alone. Middle : BundleTrack with different feature matchers. Bottom : our online pose graph optimization (PGO) module paired with different pose estimators. As visible, our module provides a significant accuracy gain in all cases.
Figure 8 : PGO runtime compared to BundleTrack on YCB Video dataset. We use the same number of correspondences for both method. Runtime is reported in millisecond per frame.
Figure 9 : Relinearization time of iSAM2. Result reports the time in milliseconds per frame averaged over frames on the YCB-Video dataset.
Method
GPU
HOI4D
YCBInEOAT
HO3D
YCB-Video
Average
One2Any
H100
30.0
30.0
30.0
30.0
30.0
FoundationPose (Tracking)
10.0
33.0
18.3
30.1
22.9
UNOPose
10.0
10.0
10.0
10.0
10.0
BundleSDF
1.4
1.3
0.9
3.5
1.8
BundleTrack(R)
1.2
1.3
0.6
2.6
1.4
Ours (RoMa v2)
H100
9.8
10.0
9.4
10.8
10.0
Table 9 : Efficiency comparison across benchmarks. Frame rate (fps) per method, reported per dataset rather than pooled, since sequence length and scene composition vary across benchmarks and can shift per-frame cost.
Method
AUC@10°
Avg.
\faCarSide
\faChair
FoundationPose(CAD)
25.5
3.6
41.1
20.6
70.5
32.2
BundleSDF
41.3
47.2
39.1
68.4
21.4
43.5
UNOPose
35.9
42.9
37.1
68.7
19.1
40.8
BundleTrack (R.)
41.0
47.4
40.5
70.4
21.5
44.2
Ours
46.6
46.5
47.0
70.0
22.2
46.5
Table 10 : Per-category rotation accuracy on five HOI4D rigid object categories. We report rotation error AUC@ 10∘ (higher is better). Our method leads on three of five non-CAD categories and on average, most notably on knife and toy car.
Figure 10 : Qualitative comparison of tracking without and with uncertainty modeling. Each pair shows the tracking result without uncertainty ( top ) and with uncertainty ( bottom ). Modeling uncertainty improves robustness in geometrically ambiguous cases, including textureless and thin objects.
Figure 11 : Failure on textureless symmetric object. Poor ADD metric due to the symmetry ambiguity.
Figure 16 : Track the object pose under motion blur.
Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion and complete occlusion due to their reliance on continuous visibility. To address these challenges, we present RRTrack, an efficient, recoverable object 6D pose tracker that enables robust tracking through fast motion and target disappearance--reappearance. RRTrack introduces a 2D--6D closed-loop tracking strategy that integrates memory-based video object segmentation (VOS) with 6D pose refinement. The 2D branch maintains target localization, and the 6D branch verifies geometric consistency before memory updates. In addition, a DINOv2-based dual-bank template matching module is developed to recover lost targets by jointly exploiting offline synthetic templates and online observation anchors while maintaining real-time efficiency. We also introduce a synthetic RGB-D benchmark comprising three robotic scenarios with fast motion and full occlusion. Experimental results on the synthetic benchmark demonstrate that RRTrack improves equal-subset mean ADD-S AR by 66.3% and ADD-S AUC by 65.7% over FoundationPose while achieving 55.2 FPS. Real-world experiments further validate the robustness of RRTrack under noisy sensing conditions. Project page: https://github.com/7kevin24/RRTrack
Junyue Li, Ye Zheng, Yifan Chen +2
College of Computer Science and Technology, Zhejiang University, Hangzhou, China · Institute of Artificial Intelligence (TeleAI), China Telecom, China · College of Future Information Technology, Fudan University, Shanghai 200433, China
Real-time 6-DoF object pose tracking is essential for many robotics applications, and several approaches exist. Yet even today's approaches remain unreliable under temporary full occlusions and rapid object motions. Once tracking is lost, most methods struggle to detect the failure and recover automatically, often requiring manual re-initialization. In this paper, we address the problem of robust model-based 6-DoF tracking of unseen objects from RGB-D data, especially in scenarios with occlusion and fast motion. We propose a novel method that combines efficient learning-based keypoint matching with optimization-based alignment and introduces a novel failure detection and recovery module. Our system monitors pose reliability, detects tracking divergence or occlusions, and performs a global re-detection and pose estimation step that robustly verifies recovery candidates before resuming tracking. Our evaluation on standard tracking benchmarks and on a new dataset of occluded and fast-moving scenes shows that our method matches state-of-the-art accuracy on easy tracking sequences, maintains high tracking speed at 57.6 frames per second, and provides the most robust tracking performance under challenging conditions. Thus, we believe that our approach is a relevant step forward in robust 6-DoF object tracking from RGB-D data.
Balázs Opra, Léo Ghafari, Thomas Stewart +1
Woven by Toyota, Inc. · University of Bonn, Germany. · University of Bonn, Center for Robotics, and the Lamarr Institute for Machine Learning and Artificial Intelligence.
Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, posed reference images, or pose annotations. Geometric foundation models provide complementary object-centric and scene-centric cues, yet SAM3D CAD is indexed by an arbitrary object-local surface parameterization, whereas reconstructed evidence is expressed in a sequence-specific world frame with partial surface coverage. To exploit this complementarity, we formulate tracking as generation-reconstruction correspondence and introduce GRC-Pose, a correspondence-based framework that combines learned correspondence prediction with robust pose estimation. Concretely, GeoCorr-Matcher estimates weighted object-scene correspondences and per-match uncertainty for each pose candidate. FGH-Solver integrates these matches through multiple robust geometric estimators and sequence-level posterior inference, while a posterior-gated memory retains only inlier-supported observations through occlusion and viewpoint change. Extensive evaluation shows that with SAM3D CAD, GRC-Pose achieves state-of-the-art Average Recall and motion retention on HOT3D, improving the latter by 58% over prior art. On classical benchmarks including YCBInEOAT and LINEMOD, it remains highly competitive.
Shiyang Liu, Weiquan Lin, Luping Xiao +4
School of Automation, Beijing Institute of Technology · Zhongguancun Academy · Artificial Intelligence Academy, Xidian University +1