Beyond Group Splits: Specimen-Level Cross-Validation and Visual Attribution for Remaining-Shelf-Life Regression in Climacteric Fruit
Authors: Rovhona Mudau, Jean Frederic Isingizwe Nturambirwe, Clement Nthambazale Nyirenda
Organizations: Department of Computer Science University of the Western Cape Cape Town, South Africa · eResearch Office University of the Western Cape Cape Town, South Africa · Department of Computer Science / eResearch Office University of the Western Cape Cape Town, South Africa
Estimating remaining shelf life (RSL) from images could provide affordable decision support for perishable produce, but evaluation protocols can substantially affect reported performance when repeated images are available from the same biological specimen. We use the Hass Avocado Ripening dataset, comprising 8,834 image-RSL pairs from 426 fruits across three storage regimes, to evaluate a frozen ImageNet-pretrained visual backbone with a lightweight regression head. Our contributions are threefold: we quantify the effect of observation-level versus specimen-disjoint evaluation, compare lightweight and heavier visual backbones under specimen-disjoint cross-validation, and examine their spatial attributions using Grad-CAM. Across ten observation-level random splits, the model achieves a mean RMSE of 2.37 days with a standard deviation of 0.03 days, whereas specimen-disjoint 5-fold cross-validation yields a mean RMSE of 3.12 days with a standard deviation of 0.11 days. The corresponding mean coefficient of determination is 0.553. A matched per-specimen comparison confirms higher error under specimen-disjoint evaluation, with a probability value below 0.001 across 426 specimens, showing that observation-level partitioning gives a substantially more optimistic estimate for this dataset and model configuration. Under specimen-disjoint evaluation, MobileNetV3-Small (0.93 million parameters) achieves accuracy comparable to ResNet-18 while providing substantially higher throughput, and Grad-CAM reveals differences in spatial attribution between the lightweight backbones. These results support specimen-disjoint evaluation and attribution analysis when assessing lightweight vision models for longitudinal shelf-life prediction.
Figures & tables
Reference
Task
Protocol category
Evaluation design
[ 13 ]
Hass avocado days-to-ripeness regression
Random sub-image split
Each hyperspectral fruit image’s sub-images were assigned to training, validation, and test sets; as a result, source-image identity was not disjoint across partitions.
[ 2 ]
Hass avocado ripening classification and shelf-life estimation
Specimen-level split
To prevent observations from the same sample from crossing partitions, avocado samples were uniquely assigned to 70/15/15 training, validation, and test subsets.
[ 14 ]
Hass avocado firmness prediction with RSL recommendation
Random image split
A total of 1,400 images were randomly shuffled into 80/10/10 training, validation, and test subsets; specimen grouping was not reported for the regression split.
[ 16 ]
Strawberry storage-age prediction
Condition-level split
Control-storage samples were used for model development and samples stored under modified-atmosphere packaging were used as the test condition.
[ 15 ]
Mango shelf-life-stage classification
Independent external validation
Internal 10-fold cross-validation was supplemented by evaluation on a separate online mango image dataset.
This study
Hass avocado RSL regression
Random, condition-level, and specimen-disjoint evaluation
Observation-level random splitting is used as a leakage diagnostic; storage regimes are held out for cross-condition analysis; primary within-dataset generalisation is assessed using specimen-disjoint 5-fold cross-validation.
TABLE I: Evaluation protocols in representative fruit ripening and shelf-life studies.
Protocol
RMSE img (d)
MAE img (d)
Rimg2
RMSE fruit (d)
MAE fruit (d)
Random observation-level split
2.37±0.03
1.73±0.03
0.744±0.006
2.42±0.04
1.76±0.02
Leave-one-storage-regime-out
5.18
4.07
-0.227
4.68
3.60
Specimen-disjoint 5-fold CV
3.12±0.11
2.35±0.08
0.553±0.023
3.07
2.30
TABLE II: Image-level and fruit-balanced performance under the three evaluation protocols.
Backbone
Par. (M)
Thr.
RMSE (d)
MAE (d)
R2
MobileNetV3-S
0.93
76/s
2.97±0.14
2.24±0.12
0.596±0.027
MobileNetV2
2.22
—
3.12±0.11
2.35±0.08
0.553±0.023
ResNet-18
11.18
11/s
3.09±0.08
2.35±0.07
0.562±0.018
TABLE III: BACKBONE COMPARISON UNDER SPECIMEN-DISJOINT 5-FOLD CROSS-VALIDATION
Stage
MobileNetV2
MobileNetV3-Small
1 (fresh)
0.44
0.12
2
0.29
0.23
3
0.16
0.28
4 (spoiled)
0.12
0.14
TABLE IV: Grad-CAM mass-on-fruit by stage
Fig. 1: Grad-CAM saliency at the ripening endpoints. MobileNetV2 (top) concentrates on the fruit when fresh and disperses once spoiled; the more accurate MobileNetV3-Small (bottom) localises to the background at both stages.