Reliable shelf monitoring is an important capability for retail automation, yet existing out-of-stock detection methods mainly operate in image space and lack metric 3D localization for downstream robotic systems. We formulate shelf monitoring as object-level 3D change detection: given two RGB-D observations captured at different times, the goal is to identify changed products and localize each change with a 3D bounding box. To support this task, we introduce ShelfChange3D, comprising 145K synthetic and 5K real-world paired RGB-D observations with object-level 3D change annotations. We further propose ChangeBox, an end-to-end framework that jointly reasons over paired observations and predicts object-level 3D change boxes. To improve localization accuracy, we introduce a geometry-based refinement stage that exploits depth and gravity prior to estimate relative pose and refine predicted boxes. Experiments show that ChangeBox outperforms existing change detection baselines, with further gains from refinement and effective transfer from synthetic to real-world observations.
Figures & tables
Fig. 1: An example of ShelfChange3D-145K, one per viewpoint-change type (rows). Red = removed, green = added.
Fig. 2: An example of a top-down view of one shelf layer. (a) The sampled layout: two SKU groups with their column/row grid cells. (b) The same layer after remove-add trajectory steps: hatched cells were vacated by removals, green rectangles are misplaced products.
Fig. 3: Camera configuration. (a) The nine nominal viewpoints on the viewing sphere around a layer anchor. (b) Azimuth/elevation of the viewpoints for one trajectory: hollow circles are nominal poses, dotted boxes the trajectory-level perturbation range, filled circles the realized trajectory-level pose, and small dots the per-frame revisit jitter over the states of the trajectory.
Fig. 4: Example of annotation pipeline for one edge. (a) Change map with difference candidates (yellow) and the one selected by the product detector (green). (b) The selected candidate and the point prompt derived from its centre (red star). (c) The SAM 3 mask. (d) The t0 RGB-D point cloud with the SAM 3D Objects mesh placed in it and the box fitted to the mesh. (e) Translucent mesh and box over t0 . (f) The 3D box over t1 , where it marks the vacated slot.
Fig. 5: Overview of the proposed method. A shared image encoder extracts patch features from the paired RGB images, while the depths are divided into patches, with each patch represented by its median depth and concatenated to the corresponding RGB feature. The resulting RGB-D tokens are fed into a change-query decoder to predict coarse 3D change boxes, which are further refined by a local box refinement module using the relative extrinsics estimated by the gravity-prior extrinsics module. Orange boxes mark removed products, blue boxes added ones.
Method
mAP 0.25 ↑
mAP 0.5 ↑
mAP 0.7 ↑
ChangeBox
87.0
57.5
14.8
Size → GT
87.2
57.7
14.8
Center → GT
97.1
95.6
94.0
Both → GT
97.4
96.7
95.5
ChangeBox+Refine
93.0
81.7
48.3
TABLE I: Error analysis of ChangeBox and evaluation of the refinement on ShelfChange3D-145K.
Fig. 6: Qualitative results. Example results for ChangeBox on ShelfChange3D-145K.
Method
mAP ↑
mAR ↑
c-err (cm) ↓
0.25
0.5
0.7
0.25
0.5
0.7
0.25
0.5
DenseChangeCap [ 29 ]
39.1
4.1
0.0
66.0
20.7
1.7
1.96
1.19
ChangeBox (R)
79.1
46.2
10.0
85.1
61.0
26.4
1.29
0.98
ChangeBox (R) + refine
88.6
74.2
36.0
92.3
81.5
53.2
0.89
0.73
ChangeBox (C)
85.9
57.3
16.3
89.5
69.2
33.3
1.14
0.91
ChangeBox (C) + refine
92.2
80.6
45.4
94.5
85.8
60.5
0.78
0.66
TABLE II: Comparison of ChangeBox variants with the baseline on ShelfChange3D-145K.
Method
Res.
mAP 0.25 ↑
mAP 0.5 ↑
mAR 0.25 ↑
c-err 0.25 ↓
ChangeBox (R)
320
61.9
23.9
72.1
1.67
ChangeBox (C)
320
75.0
37.3
82.5
1.44
ChangeBox (D)
320
73.6
35.8
82.0
1.44
ChangeBox (R)
640
79.1
46.2
85.1
1.29
ChangeBox (C)
640
85.9
57.3
89.5
1.14
ChangeBox (D)
640
87.0
57.5
90.5
1.14
TABLE III: Evaluation of different visual representation and resolution on ShelfChange3D-145K.
Method
Depth
mAP 0.25 ↑
mAP 0.5 ↑
mAR 0.25 ↑
c-err 0.25 ↓
ChangeBox (R)
✗
72.4
38.4
81.0
1.36
ChangeBox (R)
✓
79.1
46.2
85.1
1.29
ChangeBox (C)
✗
83.0
52.3
88.1
1.16
ChangeBox (C)
✓
85.9
57.3
89.5
1.14
ChangeBox (D)
✗
80.5
45.1
86.7
1.28
ChangeBox (D)
✓
87.0
57.5
90.5
1.14
TABLE IV: Evaluation of effect of the depth information on ShelfChange3D-145K.
Fig. 7: Accuracy and center error vs. relative camera rotation between t0 and t1 . The <5∘ bin holds 61% of the test pairs.
Fig. 8: Extrinsics estimation on the test set. Left: solve rate per view type. Right: wrongly accepted per view type, both with and without the forward-backward check.
Training
Metrics
Syn. pre-train
Real
Refine
mAP 0.25 ↑
mAP 0.5 ↑
mAR 0.25 ↑
c-err 0.25 ↓
✗
✓
✗
21.0
0.8
44.2
3.34
✓
✗
✗
46.7
2.2
65.3
2.87
✓
✓
✗
92.9
37.5
95.6
2.20
✓
✓
✓
97.9
50.1
98.5
1.95
TABLE V: Evaluation of Sim2Real transfer on ShelfChange3D-5K.
Fig. 9: Restocking Robot. For each demonstration, the top row shows the head-camera views: the fully stocked shelf at t0 , the shelf at t1 after products were removed, the predicted boxes of the removed products, and the shelf after restocking. Subsequent panels show robots demonstration.
Method
Schedule
mAP 0.25 ↑
mAP 0.5 ↑
mAR 0.25 ↑
c-err 0.25 ↓
ChangeBox (R)
FT
66.5
28.5
77.4
1.56
ChangeBox (R)
frozen
64.1
24.7
74.3
1.66
ChangeBox (R)
frozen → FT
79.1
46.2
85.1
1.29
ChangeBox (C)
FT †
60.4
17.8
74.2
1.86
ChangeBox (C)
frozen
72.9
33.7
80.8
1.51
ChangeBox (C)
frozen → FT
85.9
57.3
89.5
1.14
TABLE VI: Evaluation of different training schedules on the ShelfChange3D-145K.
Fig. 10: Qualitative results. Example results for ChangeBox on ShelfChange3D-5K.
Fig. 11: One remove-add trajectory seen from the front camera. Dashed red boxes mark the products removed at that step at their previous position, solid green boxes mark the misplaced products added at that step.
Dataset
Domain
Task
Scale
Annotation
Change detection in image space
VL-CMU-CD [ 18 ]
street
segmentation
1,362 pairs
2D mask
ChangeSim [ 27 ]
industrial
segmentation
130K images
2D mask
PSCD [ 19 ]
street
segmentation
770 pairs
2D mask
3D and multiview scene change detection
3RScan [ 28 ]
indoor rooms
re-localization
1,482 scans
instance pose
TABLE VII: Comparison with existing change detection datasets and retail datasets.
Fig. 13: Examples from the ShelfChange3D-5K test split, t0 on the left and t1 on the right.
Recovering the 9D pose of objects, both their 6D pose and 3D dimensions, under clutter and occlusion is a core requirement for warehouse automation, logistics, and manufacturing. Model-based methods are accurate but assume an instance-specific CAD model for every object, which is costly to maintain as inventories change. Model-free and category-level methods relax this assumption, yet they remain vulnerable to the symmetry, weak texture, and heavy occlusion that characterize stacked storage boxes, and they ignore the strong structural priors such scenes provide. We present \textbf{AnyBox}, an efficient zero-shot framework that exploits the geometric regularity of boxes to jointly recover pose and dimensions from a single RGB-D observation. Starting from a canonical category template, AnyBox alternates between pose and scale estimation, using the discrepancy between the reprojected template and the observed mask to drive a binary search over box dimensions. Two lightweight components make this practical: a depth-consistency filter that rejects the implausible hypotheses induced by box symmetry, and an early-stopping rule that replaces the remaining search with a single closed-form update. On public benchmarks and an in-house warehouse dataset, AnyBox improves detection AP by up to 36 points, more than doubling the previous best, and approaches instance-level pipelines that have access to ground-truth CAD models. These gains transfer downstream, raising success by 28% on a cluttered robotic box-shelving task.
Yintao Ma, Sajjad Pakdamansavoji, Charles Eret +5
1Huawei Technologies Canada · University of Waterloo
Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector's predicted depth degrades accuracy, even falling below the detector's own prediction. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM's vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.
Zhihao Zhang, Gengwei Zhang, Tianlong Chen +1
Michigan State University · University of North Carolina at Chapel Hill
3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods predominantly operate in the 2D domain, where viewpoint variations are mistaken for physical changes and depth is unavailable. While visual geometry foundation models like VGGT rapidly produce dense point clouds from unposed images, independent per-epoch reconstruction encounters fundamental obstacles: unpredictable inter-epoch scale ambiguity, registration-change paradox where scene changes corrupt alignment, and pervasive edge-flying noise. To address these challenges, we present VGGT-CD, a training-free pipeline decoupling cross-temporal registration from dynamic-change interference. In the Coarse Stage, sparse keyframe joint inference establishes a unified metric space and yields an initial Sim(3) prior. In the Fine Stage, dense reconstructions are purified by isolating static-background correspondences. A closed-form centroid alignment refines the translation while locking scale and rotation, using a residual self-check to mathematically guarantee non-degradation. Evaluated on an 11-scene benchmark from the World Across Time dataset, VGGT-CD reduces Absolute Trajectory Error by 44% outdoors and 59% indoors. It completes registration over 6 times faster, producing high-purity 3D change maps without task-specific training.
Wei Zhang, Songhua Li, Yihang Wu +2
Northwestern Polytechnical University, Xi'an, China · Northwestern Polytechnical University Xian, 710072, China