Product Identification has sprung up to become one of the most challenging problems in the automation of the retail industry. With the new industry 5.0 standards, automated inventory management, and catalog creation tasks are vitally important. Object identification models have emerged as a viable answer with their unprecedented identification and localization accuracy. However, the close-knit rack design of supermarkets generates the problem of angle variation in capturing images. The angle-variant densely packed images(a single image contains many objects) become overwhelming for these models alone. In this paper, we try to supplement object detection models with traditional Hough transform (HT) and homogeneous estimation concepts. We study the effect of rectified images using homography estimation and hough transform and their limitations on the problem of grocery identification. We make a case for creating a new dataset to test the effects of such rectification and produce analytical results on different scenarios of angle variation and object densities per image. Extensive experiments on different object detection models suggest that image rectification of angled images improves the detection accuracy of grocery products in images. The results also highlight the limitation of rectification on the angle of image capture and the object density of the image.
Figures & tables
Figure 1: The illustration depicts the scheme of grocery identification. A user captures an image of the grocery shelf. The image is then passed through an object detection module that identifies the input image’s products. The identified products are then passed through a description database that associates product descriptions with each product.
Figure 2: The figure shows the process of rectifying angled images. As input, two pairs of images (Right angled+front view, left angled+front view) are taken as input, and appropriate homography is estimated. Hough lines are generated on the test image, which is associated the image’s alignment. Input image with appropriate homography generates a rectified image.
Dataset
DS
MS
GD
BC
FG
RW
Grozi-120
×
×
×
✓
✓
✓
Freiburg
×
×
×
✓
✓
✓
RPC
✓
×
×
✓
✓
✓
SKU110K
✓
✓
✓
×
×
✓
Grocer-Help
✓
✓
✓
✓
✓
✓
Table 1: Comparison of representative grocery datasets with Grocer-Help.
Figure 3: The image illustrates the procedure for generating Hough Lines and Image Instances of Hough Transform on left-aligned and right-aligned images
Figure 4: The figure illustrates the effect of Image rectification on Left and Right Aligned Images using respective aligned Homography
Dataset
Images
Classes
YOLOv3
YOLOv5m
YOLOv7
YOLOv8m
YOLOv9
RetinaNet
GroZi-120 [ 24 ]
11870
120
24.1
23.9
25.61
25.2
26.3
24.1
Webmarket [ 25 ]
3153
200
29.4
29.2
31.2
31.8
32.1
31.3
GroZi-3.2K [ 26 ]
8350
80
27.1
26.1
27.5
26.7
29.1
26.6
FreiBurg [ 21 ]
4947
25
88.1
73
79.35
92.4
93.5
78.85
SKU110K [ 20 ]
11762
1
91.7
89
88.49
90.9
92.2
90.0
Table 2: mAP of SOTA models on existing Grocery Dataset
Figure 5: The Figure highlights the camera angle’s effect on the image’s view. Image (a) showcases a frontal view of a grocery rack, (b) Depicts the view of the rack when the image is taken from a left-aligned camera, and (c) Depicts the view of the rack when the image is taken from a right aligned camera.
Figure 6: The Figure highlights the Precision-Recall curve of best-performing algorithms on our dataset
Model
FLOPS
Parameters
Non Rectified Images
Rectified Images
Frontal View
Right Angled
Left Angled
mAP 0.5
mAP 0.9
mAP 0.5
mAP 0.9
mAP 0.5
mAP 0.9
mAP 0.5
mAP 0.9
YOLO v3
283.7 G
103.83 M
87.6
73.8
52.23
38.67
51.1
39.7
54.79
40.12
YOLO v5n
8.8 G
2.86 M
60.1
46.2
49.12
40.1
49.4
41.38
56.2
45.1
YOLO v5s
24.4 G
9.19 M
83.9
65.5
52.3
41.2
53.1
41.7
57.2
45.2
YOLO v5m
64.9 G
25.17 M
88.3
71.4
54.1
41.6
55.1
45.2
58.6
43.9
Table 3: The table summarises the accuracy of different Object detection models on our dataset with a frontal view, angled view, and rectified images with mAP 0.5 and mAP 0.9 as the evaluation metric (Here G stands for Giga and M stands for Million)
Angle
O/I ≤ 10
O/I ≥ 11
mAP 0.5
mAP 0.5 (R.I)
mAP 0.5
mAP 0.5 (R.I)
10 - 30
41.2
45.5
14.3
13.9
30 - 60
74.8
80.2
34.2
36.1
60 - 80
83.5
84.0
56.2
57.1
110 - 130
83.1
83.2
57.1
56.3
130 - 150
71.6
75.7
36.1
36.8
Table 4: The table highlights the effect of angle variation in mAP 0.5 on densely packed and sparsely packed images. Here O/I is Objects per Image and R.I is Rectified Image
Figure 7: The Figure highlights the gain perceived in all the existing algorithms when run with rectified images in comparison to the angled images on Grocer-Help dataset
O/I
30 ≤ right_alignment ≤ 60
30 ≤ left_alignment ≤ 60
mAP 0.5
mAP 0.5 (R.I)
mAP 0.5
mAP 0.5 (R.I)
less than 5
89.8
89.9
88.4
89.0
5 - 10
85.3
86.4
84.1
84.9
10 - 20
79.2
83.1
78.8
83.1
20 - 40
49.1
50.0
48.2
50.0
more than 40
18.2
15.1
19.1
18.2
Table 5: The table highlights the effect of Objects per Image (O/I) with the viewing angle restricted to 30 ≤θ≤ 60. R.IstandforRectifiedImage
Automated product recognition is a cornerstone of modern retail intelligence; however, accurately matching real-world, in-store images against extensive corporate catalogs remains a major scalability bottleneck for large-scale applications. In this work, we address this challenge by reformulating the task as an embedding-based cross-domain retrieval problem rather than a standard closed-set classification task. Specifically, we define the objective as retrieving the most corresponding catalog reference image for a given real-world product query crop from an expansive inventory. To bridge the severe domain gap between pristine studio packshots and noisy in-store queries, we introduce a novel catalog-to-real multi-stage contrastive learning paradigm (Cat2Real). This framework fine-tunes a vision backbone by systematically exploiting both item-level and image-level similarities to drive targeted hard negative mining. Extensive empirical evaluations demonstrate that our paradigm scales seamlessly to unseen products and categories, yielding outstanding zero-shot generalization performance even in the complete absence of real-world training images for novel inventory.
Multimodal product retrieval (MPR) underpins checkout-free retail and automated inventory systems, yet it demands fine-grained SKU discrimination that standard vision-language benchmarks fail to capture. We present the first systematic zero-shot evaluation of 190 open-source VLMs on the MPR task of the GroceryVision Challenge, isolating pre-training data, architecture, and input resolution. Our analysis yields three actionable findings. \textbf{(1) Data quality trumps scale.} Switching from raw web-scrapes to filtered datasets delivers up to 16.6% accuracy gains, exceeding the benefit of doubling model parameters. \textbf{(2) Efficient models can win.} MobileCLIP-B (150M parameters) outperforms 351M counterparts trained on noisy data. We introduce \textit{semantic power density} (φ), an efficiency metric that penalizes sub-threshold accuracy. \textbf{(3) A precision gap persists.} State-of-the-art models achieve 94.5% Recall@5 but suffer a 17.5% drop at Recall@1, revealing that contrastive embeddings cluster categories effectively but fail to rank visually similar SKUs. Code and evaluation scripts are available at https://github.com/upeee/openmpr.
Emmanuel G. Maminta, Rowel O. Atienza
AI Graduate Program, University of the Philippines, Diliman, Quezon City · EEEI, University of the Philippines, Diliman, Quezon City
Image correction and rectangling are valuable tasks in practical photography systems such as smartphones. Recent remarkable advancements in deep learning have undeniably brought about substantial performance improvements in these fields. Nevertheless, existing methods mainly rely on task-specific architectures. This significantly restricts their generalization ability and effective application across a wide range of different tasks. In this paper, we introduce the Unified Rectification Framework (UniRect), a comprehensive approach that addresses these practical tasks from a consistent distortion rectification perspective. Our approach incorporates various task-specific inverse problems into a general distortion model by simulating different types of lenses. To handle diverse distortions, UniRect adopts one task-agnostic rectification framework with a dual-component structure: a {Deformation Module}, which utilizes a novel Residual Progressive Thin-Plate Spline (RP-TPS) model to address complex geometric deformations, and a subsequent Restoration Module, which employs Residual Mamba Blocks (RMBs) to counteract the degradation caused by the deformation process and enhance the fidelity of the output image. Moreover, a Sparse Mixture-of-Experts (SMoEs) structure is designed to circumvent heavy task competition in multi-task learning due to varying distortions. Extensive experiments demonstrate that our models have achieved state-of-the-art performance compared with other up-to-date methods.
Linwei Qiu, Gongzhe Li, Xiaozhe Zhang +2
Tianmushan Laboratory, Beihang University, Hangzhou, China · State Key Laboratory of High-Efficiency Reusable Aerospace Transportation Technology, Beijing, China · School of Astronautics, Beihang University, Beijing, China +2