Organizations: Dept. of Information Systems, University of Maryland Baltimore County, Baltimore, MD 21250, USA · Dept. of Electrical Engineering, Indian Institute of Technology Kharagpur, West Bengal 721302, India · Dept of Data Science and Engineering, Indian Institute of Science Education and Research Bhopal 462066, India · UJM-Saint-Etienne, CNRS, Institut d’Optique Graduate School, Laboratoire Hubert Curien, Saint-Etienne, France
The same fruit appears in a bunch, unpicked, peeled, bagged in plastic, or sliced on a dish, so automated fruit classification in the wild (AFCW) must absorb wide intra- class and narrow inter-class variability in shape, size, colour and texture. Convolutional networks route information through pooling, which discards the pose and location of the region of interest and therefore generalises poorly across these presentations. We propose FruitCapsNet, a capsule network whose Fruit Capsules replace the standard convolutional front end with dilated convolutions: the receptive field grows exponentially at constant parameter cost, so each capsule encodes multi-scale context before dynamic routing resolves part whole spatial agreement. Hyper-parameters, including the dilation factor, are selected by Bayesian optimisation rather than grid search. On three public datasets (SMP, FruitsGB, Fruits-360) and a new 19-class, 10,639-image in-the-wild dataset (PD-19), FruitCapsNet exceeds ten fine-tuned transfer-learning backbones at one-third the depth, with the largest margin (+2.7% over the nearest competitor) on the hardest set. Grad-CAM saliency propagated from the DigitCaps layer shows that the improvement comes from attributing decisions to whole-fruit regions rather than to object edges, giving post-hoc evidence that the gain is not a dataset artefact.
Figures & tables
Figure 1: Block diagram of the FruitCapsNet-based AFCW.
Figure 2: Modular schematic of FruitCapsNet. Three dilated convolutional layers feed PrimaryCaps; DigitCaps emits one capsule per class; the decoder reconstructs the input as a regulariser.
Figure 3: Receptive field for dilation (a) d=1 , (b) d=2 , (c) d=3 .
Dataset
Cls.
Images
Res.
Character
SMP [ 61 ]
15
2,633
1024×768
bagged, shadowed, multi-count
FruitsGB [ 62 ]
12
12,000
256×256
6 fruits × good/bad quality
Fruits-360 [ 64 ]
81
55,244
100×100
pre-segmented, white bg.
PD-19 (ours)
19
10,639
variable
in the wild, multi-class/image
Table I: Datasets. Per-class breakdowns in supplementary material.
Figure 4: SMP samples: illumination, pose and count vary within class.
Figure 5: PD-19 samples: multiple categories per image, cluttered backgrounds, occlusion, and cut/plated/unpicked presentations.
Figure 6: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the SMP dataset.
Figure 7: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the FruitsGB dataset.
Figure 8: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the Fruit 360 dataset.
Figure 9: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the PD-19 dataset.
Figure 12: PD-19: inputs ( 1st row), dense-layer maps ( 2nd ), Grad-CAM for FruitCapsNet ( 3rd ) and DenseNet ( 4th ). FruitCapsNet retains whole-object attribution under occlusion and clutter.
Dataset
Acc.
F1
Rec.
Prec.
κ
FruitsGB
98.84
98.54
98.98
98.17
98.72
SMP
99.19
99.39
99.02
99.49
99.12
Fruits-360
99.31
99.28
99.41
99.07
99.28
PD-19
98.28
98.25
98.55
97.91
98.43
Table II: FruitCapsNet test performance (%), mean over repeated splits.
Network
Depth
PM
SMP
FruitsGB
Fruits-360
PD-19
ResNet-50 [ 65 ]
50
25.63
49.86
72.02
80.28
89.01
ResNet-101 [ 65 ]
101
44.70
90.85
97.16
97.20
93.05
ResNet-152 [ 65 ]
152
60.38
92.18
98.05
96.96
93.51
VGG-16 [ 66 ]
16
138.3
96.91
97.77
98.00
93.84
VGG-19 [ 66 ]
19
143.6
97.46
96.97
97.58
91.22
DenseNet-201 [ 67 ]
201
20.0
98.39
98.52
98.41
95.61
Table III: Comparison with fine-tuned pre-trained backbones (accuracy, %).
Set
Ref
Method
Acc.
SMP
[ 73 ]
colour/shape/texture + wavelet co-occurrence
86.00
[ 74 ]
autocorrelogram, CCV, BIC, LAS + ML fusion
98.80
[ 75 ]
LBP, HOG, GaborLBP, CNN+SVM, R-CNN
98.50
[ 48 ]
CNN, VGG-16
88.35
ours
FruitCapsNet + BO
99.19
Fruits -360
[ 76 ]
pure CNN, global average pooling
98.88
Table IV: Comparison with published results. Prior art on SMP and Fruits-360 spans hand-crafted descriptors and CNN detectors; FruitsGB and PD-19 have a single deep baseline.
Configuration
SMP
FruitsGB
Fruits-360
PD-19
CapsNet [ 24 ] , d=1 , hand-tuned
98.32
97.46
98.54
96.72
+ dilation d=2
98.81
98.15
99.02
97.81
+ dilation d=3
98.56
97.92
98.83
97.35
+ dilation d=2 + BO (full)
99.19
98.84
99.31
98.28
Table V: Ablation (accuracy, %; mean over 3 seeds).
1Efrei Research Lab, Paris Panthéon-Assas University, Paris, France · Institute of Applied Computer Science, Lodz University of Technology, Łódź, Poland
Engineering Department, University of Palermo, V.le delle Scienze, Ed. 6, Palermo, 90128, Italy. · Computer Vision Lab, Delft University of Technology, Van Mourik Broekmanweg 6,2026 Delft, 2628 XE, The Netherlands.