Organizations: Dept. of Information Systems, University of Maryland Baltimore County, Baltimore, MD 21250, USA · Dept. of Electrical Engineering, Indian Institute of Technology Kharagpur, West Bengal 721302, India · Dept of Data Science and Engineering, Indian Institute of Science Education and Research Bhopal 462066, India · UJM-Saint-Etienne, CNRS, Institut d’Optique Graduate School, Laboratoire Hubert Curien, Saint-Etienne, France
The same fruit appears in a bunch, unpicked, peeled, bagged in plastic, or sliced on a dish, so automated fruit classification in the wild (AFCW) must absorb wide intra- class and narrow inter-class variability in shape, size, colour and texture. Convolutional networks route information through pooling, which discards the pose and location of the region of interest and therefore generalises poorly across these presentations. We propose FruitCapsNet, a capsule network whose Fruit Capsules replace the standard convolutional front end with dilated convolutions: the receptive field grows exponentially at constant parameter cost, so each capsule encodes multi-scale context before dynamic routing resolves part whole spatial agreement. Hyper-parameters, including the dilation factor, are selected by Bayesian optimisation rather than grid search. On three public datasets (SMP, FruitsGB, Fruits-360) and a new 19-class, 10,639-image in-the-wild dataset (PD-19), FruitCapsNet exceeds ten fine-tuned transfer-learning backbones at one-third the depth, with the largest margin (+2.7% over the nearest competitor) on the hardest set. Grad-CAM saliency propagated from the DigitCaps layer shows that the improvement comes from attributing decisions to whole-fruit regions rather than to object edges, giving post-hoc evidence that the gain is not a dataset artefact.
Figures & tables
Figure 1: Block diagram of the FruitCapsNet-based AFCW.
Figure 2: Modular schematic of FruitCapsNet. Three dilated convolutional layers feed PrimaryCaps; DigitCaps emits one capsule per class; the decoder reconstructs the input as a regulariser.
Figure 3: Receptive field for dilation (a) d=1 , (b) d=2 , (c) d=3 .
Dataset
Cls.
Images
Res.
Character
SMP [ 61 ]
15
2,633
1024×768
bagged, shadowed, multi-count
FruitsGB [ 62 ]
12
12,000
256×256
6 fruits × good/bad quality
Fruits-360 [ 64 ]
81
55,244
100×100
pre-segmented, white bg.
PD-19 (ours)
19
10,639
variable
in the wild, multi-class/image
Table I: Datasets. Per-class breakdowns in supplementary material.
Figure 4: SMP samples: illumination, pose and count vary within class.
Figure 5: PD-19 samples: multiple categories per image, cluttered backgrounds, occlusion, and cut/plated/unpicked presentations.
Figure 6: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the SMP dataset.
Figure 7: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the FruitsGB dataset.
Figure 8: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the Fruit 360 dataset.
Figure 9: Over the epochs, variation in FruitCapsNet‘s (a) accuracy and total loss, (b) different model loss components (encoder & decoder loss), and (c) comparison of network convergence for different TL architectures for the PD-19 dataset.
Figure 12: PD-19: inputs ( 1st row), dense-layer maps ( 2nd ), Grad-CAM for FruitCapsNet ( 3rd ) and DenseNet ( 4th ). FruitCapsNet retains whole-object attribution under occlusion and clutter.
Dataset
Acc.
F1
Rec.
Prec.
κ
FruitsGB
98.84
98.54
98.98
98.17
98.72
SMP
99.19
99.39
99.02
99.49
99.12
Fruits-360
99.31
99.28
99.41
99.07
99.28
PD-19
98.28
98.25
98.55
97.91
98.43
Table II: FruitCapsNet test performance (%), mean over repeated splits.
Network
Depth
PM
SMP
FruitsGB
Fruits-360
PD-19
ResNet-50 [ 65 ]
50
25.63
49.86
72.02
80.28
89.01
ResNet-101 [ 65 ]
101
44.70
90.85
97.16
97.20
93.05
ResNet-152 [ 65 ]
152
60.38
92.18
98.05
96.96
93.51
VGG-16 [ 66 ]
16
138.3
96.91
97.77
98.00
93.84
VGG-19 [ 66 ]
19
143.6
97.46
96.97
97.58
91.22
DenseNet-201 [ 67 ]
201
20.0
98.39
98.52
98.41
95.61
Table III: Comparison with fine-tuned pre-trained backbones (accuracy, %).
Set
Ref
Method
Acc.
SMP
[ 73 ]
colour/shape/texture + wavelet co-occurrence
86.00
[ 74 ]
autocorrelogram, CCV, BIC, LAS + ML fusion
98.80
[ 75 ]
LBP, HOG, GaborLBP, CNN+SVM, R-CNN
98.50
[ 48 ]
CNN, VGG-16
88.35
ours
FruitCapsNet + BO
99.19
Fruits -360
[ 76 ]
pure CNN, global average pooling
98.88
Table IV: Comparison with published results. Prior art on SMP and Fruits-360 spans hand-crafted descriptors and CNN detectors; FruitsGB and PD-19 have a single deep baseline.
Configuration
SMP
FruitsGB
Fruits-360
PD-19
CapsNet [ 24 ] , d=1 , hand-tuned
98.32
97.46
98.54
96.72
+ dilation d=2
98.81
98.15
99.02
97.81
+ dilation d=3
98.56
97.92
98.83
97.35
+ dilation d=2 + BO (full)
99.19
98.84
99.31
98.28
Table V: Ablation (accuracy, %; mean over 3 seeds).
Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-ranging advantages in agriculture field for both pre-harvest and post-harvest management. Accurate and timely FRP can be achieved using machine/deep learning-based hyperspectral image classification techniques. However, challenges including the limited availability of labeled data and the lack of robust methods generalizable to various hyperspectral cameras and fruit types can compromise the effectiveness of hyperspectral image-based FRP. Addressing these challenges, this paper introduces Fruit-HSNet, a machine learning architecture specifically designed for hyperspectral classification of fruit ripeness. Fruit-HSNet incorporates a spatio-spectral feature extraction module based on Fourier Transform and central pixel spectral signature followed by learnable feature fusion and a classifier optimized for ripeness classification. The proposed architecture was evaluated using the DeepHS Fruit dataset, the largest publicly available labeled real-world hyperspectral dataset for predicting fruit ripeness, which includes five different types of fruits-avocado, kiwi, mango, kaki, and papaya-captured with three distinct hyperspectral cameras at various stages of ripeness. Experimental results highlight that Fruit-HSNet substantially outperforms existing deep learning methods, from baseline to state-of-the-art models, with improvements of 12%, achieving a new state-of-the-art overall accuracy of 70.73%.
Ahmed Baha Ben Jmaa, Faten Chaieb, Anna Fabijańska
1Efrei Research Lab, Paris Panthéon-Assas University, Paris, France · Institute of Applied Computer Science, Lodz University of Technology, Łódź, Poland
Fine-grained fruit classification is a critical yet challenging task in agricultural computer vision, primarily hindered by a severe shortage of high-quality datasets and the high visual similarity between classes. To address these challenges, we first constructed a comprehensive dataset comprising 306 fruit categories with 116,233 samples. Moreover, we propose FruitEnsemble, a practical two-stage dynamic inference framework designed to overcome the generalization limitations of static single-model architectures. In the first stage, FruitEnsemble employs a validation-calibrated weighted ensemble of heterogeneous backbones to generate a robust Top-3 candidate pool. To tackle difficult samples, we introduce an expert arbitration mechanism: when ensemble confidence falls below 0.6, a multimodal large language model (MLLM) is triggered to perform rigorous visual verification by integrating external botanical descriptions using Chain-of-Thought (CoT) reasoning. Furthermore, we optimized the training pipeline with a hard sample-aware joint loss. Extensive experiments demonstrate that FruitEnsemble achieves a classification accuracy of 70.49% and outperforms existing state-of-the-art models. Our framework provides an efficient, deployment-oriented solution for real-world agricultural visual sorting and quality inspection tasks.
Enhui Yu, Junhui Li, Ruitong Lu +2
University of Science and Technology Liaoning · Chuzhou University · Yeshiva University
This paper proposes a new method to improve the training efficiency of deep convolutional neural networks. During training, the method evaluates scores to measure how much each layer's parameters change and whether the layer will continue learning or not. Based on these scores, the network is scaled down such that the number of parameters to be learned is reduced, yielding a speed up in training. Unlike state-of-the-art methods that try to compress the network to be used in the inference phase or to limit the number of operations performed in the backpropagation phase, the proposed method is novel in that it focuses on reducing the number of operations performed by the network in the forward propagation during training. The proposed training strategy has been validated on two widely used architecture families: VGG and ResNet. Experiments on MNIST, CIFAR-10 and Imagenette show that, with the proposed method, the training time of the models is more than halved without significantly impacting accuracy. The FLOPs reduction in the forward propagation during training ranges from 17.83% for VGG-11 to 83.74% for ResNet-152. These results demonstrate the effectiveness of the proposed technique in speeding up learning of CNNs. The technique will be especially useful in applications where fine-tuning or online training of convolutional models is required, for instance because data arrive sequentially.
Giorgio Cruciata, Luca Cruciata, Liliana Lo Presti +2
Engineering Department, University of Palermo, V.le delle Scienze, Ed. 6, Palermo, 90128, Italy. · Computer Vision Lab, Delft University of Technology, Van Mourik Broekmanweg 6,2026 Delft, 2628 XE, The Netherlands.