A Data-Centric Review of Plant Disease Datasets: Taxonomy, Critical Analysis, Environmental Variability, and Implications for Precision Agriculture
Organizations: Department of Information Technology, National Institute of Technology, Hazratbal Srinagar 190006, Jammu & Kashmir, India · Department of Computer Science and Engineering, Indian Institute of Technology, Ropar Rupnagar 140001, Punjab, India
Abstract
Despite rapid advances in artificial intelligence, reliable real-world plant disease detection remains a persistent challenge. Visual and deep learning approaches have shown promising results, but their deployment under field conditions remains limited. A key bottleneck is the reliance on laboratory-generated datasets that lack environmental diversity, realistic backgrounds, and balanced class distributions, resulting in poor generalization. In contrast, datasets collected directly from agricultural environments capture natural variability and better reflect challenges faced by farmers across regions. This review presents a critical analysis of visual and deep learning approaches for plant disease detection, with emphasis on plant disease datasets. It establishes a taxonomy based on acquisition setting, accessibility, plant diversity, disease composition, class structure, and imbalance severity, and examines their implications for model generalization and real-world deployment. A comparative analysis of laboratory and real-field datasets identifies critical gaps that hinder disease detection. The review further analyzes how multi-level dataset imbalance, including intra-class, inter-crop, and cross-dataset imbalance, and limited environmental variability affect model performance and robustness, an area insufficiently examined in existing surveys. Beyond image-based approaches, it highlights the importance of integrating environmental parameters such as temperature, humidity, and leaf wetness with image data to improve prediction under dynamic field conditions. Finally, the review identifies key challenges, research gaps, and future directions concerning dataset construction, environmental variability, structural imbalance, standardization, and multimodal disease monitoring. It provides a foundation for developing next-generation multimodal frameworks for precision agriculture.
Figures & tables
| Ref | Year | Model Coverage | Architecture Analysis | Dataset Inclusion | Dataset Analysis | Data Diversity | Class Imbalance | Acquisition Setting | Dataset Availability | Env Consideration | Dataset Limitations |
| [ 50 ] | 2025 | ✓ | Moderate | ✓ | Limited | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| [ 53 ] | 2025 | ✓ | Moderate | ✓ | Limited | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| [ 51 ] | 2025 | ✓ | Comprehensive | ✓ | Limited | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| [ 52 ] | 2024 | ✓ | Comprehensive | ✓ | Limited | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| [ 9 ] | 2024 | ✓ | Moderate | ✓ | Limited | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| [ 54 ] | 2024 | ✓ | Moderate | ✓ | Limited | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Database | Retrieved Articles |
| ACM Digital Library | 50 |
| Scopus | 1654 |
| SpringerLink | 1328 |
| IEEE Xplore | 1742 |
| Google Scholar | 1934 |
| Other sources | 76 |
| Considered articles | Unconsidered articles |
| Studies on plant disease detection using deep learning techniques. | Studies on general agriculture without focus on disease detection or deep learning applications |
| Articles that describe or analyze plant disease datasets relevant to deep learning models. | Articles focused on traditional plant pathology methods without using datasets or DL models |
| Research that includes single-plant or | |
| multi-plant datasets for identifying diseases in crops or plants | Studies that do not provide or analyze datasets or only mention deep learning models in theoretical terms. |
| Articles discussing image datasets specifically collected for plant disease classification | Research on crop yield prediction or growth monitoring without a focus on disease detection. |
| Articles discussing the comparison of DL and datasets in plant disease detection across different plant species. | Papers using non-disease-related datasets (e.g., plant growth data, soil composition data) without disease-specific context. |
| Ref | Dataset | Year | Acquisition | Availability | No. of Crops | No. of Classes | No. of Images |
| Single-Plant Datasets | |||||||
| [ 61 ] | ALDD Apple Dataset | 2019 | Lab | Private | 1 | 5 | 26,377 |
| [ 62 ] | Soybean Dataset | 2019 | Lab | Private | 1 | 7 | 1,470 |
| [ 63 ] | Citrus Dataset | 2022 | Lab | Private | 1 | 3 | 2,684 |
| [ 64 ] | Rice Dataset | 2022 | Lab | Private | 1 | 3 | 442 |
| [ 65 ] | Apple Plant Dataset | 2022 | Lab | Private | 1 | 10 | 5,201 |
| Ref | Dataset Name | Access Link |
| Publicly Available Single-Plant Datasets | ||
| [ 69 ] | Rice Dataset | https://archive.ics.uci.edu/ml/datasets/Rice+Leaf+Diseases |
| [ 70 ] | Soybean Dataset | https://datadryad.org/stash/dataset/doi:10.5061/dryad.41ns1rnj3 |
| [ 45 ] | Beans Dataset | https://github.com/AI-Lab-Makerere/ibean/tree/master |
| [ 20 ] | Citrus Plant Dataset | https://www.kaggle.com/datasets/drlisbeek/citrus-leaves-prepared |
| Publicly Available Multi-Plant Datasets | ||
| Dataset | Class of diseases | No. of images | Total Images | Skewness measure |
|---|---|---|---|---|
| Apple dataset ALDD [ 61 ] | Alternaria Leaf Spot | 5343 | 26,377 | 0.155 |
| Apple-Rust | 5694 | |||
| Mosaic-Virus | 4875 | |||
| Gray-Spot | 4810 | |||
| Brown-Spot | 5655 | |||
| Soybean Dataset [ 62 ] | Bacterial | 200 | 1,470 | 0.294 |
| Plant Name | Healthy | Diseased | Total Images (Crop) | % of Dataset | Dominant Class | Skewness Measure |
| Arjun | 220 | 232 | 452 | 11.08% | Diseased | 0.052 |
| Mango | 170 | 265 | 435 | 10.66% | Diseased | 0.358 |
| Bael | 0 | 118 | 118 | 2.89% | Diseased | 1.00 |
| Guava | 277 | 142 | 419 | 10.27% | Healthy | 0.487 |
| Chinar | 103 | 120 | 223 | 5.46% | Diseased | 0.142 |
| Lemon | 159 | 77 | 236 | 5.78% | Healthy | 0.516 |
| Plant Name | Class of Diseases | No. of Images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Apple | Apple Cider Rust | 276 | 3,172 | 5.84% | Healthy | 0.833 |
| Apple Black Rot | 621 | |||||
| Apple Scab | 630 | |||||
| Apple Healthy Leaf | 1645 | |||||
| Peach | Peach Bacterial Spot | 2297 | 2,657 | 4.89% | Bacterial Spot | 0.843 |
| Peach Healthy Leaf | 360 |
| Plant Name | Class of diseases | No of images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness measure |
| Apple | Apple Scab Leaf | 92 | 272 | 10.59% | Apple Scab | 0.33 |
| Apple Rust Leaf | 89 | |||||
| Apple Healthy Leaf | 91 | |||||
| Blueberry | Blueberry Healthy Leaf | 117 | 117 | 4.55% | Healthy Leaf | 1.00 |
| Blueberry Diseased Leaf | 0 | |||||
| Cherry | Cherry Healthy Leaf | 57 | 57 | 2.21% | Healthy Leaf | 1.00 |
| Plant Name | Class of Disease | No. of Images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Apple | Apple Aphis spp | 162 | 1416 | 31.84% | Venturia inaequalis | 0.744 |
| Apple Venturia inaequalis | 633 | |||||
| Apple Monilia lava | 255 | |||||
| Apple Eriosoma lanigerum | 366 | |||||
| Peach | Peach Monilia lava | 314 | 746 | 16.77% | Parthenolecanium corni | 0.264 |
| Peach Parthenolecanium corni | 427 |
| Plant Name | Class of Disease | No. of Images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Corn | Leaf Blight | 1655 | 3,194 | 37.0% | Leaf Blight | 0.996 |
| Brown Spots | 220 | |||||
| Corn Stripe | 164 | |||||
| Corn Yellowing | 493 | |||||
| Corn Streak | 251 | |||||
| Chlorotic Leaf Spot | 26 |
| Plant Name | Class of Disease | No. of Images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Apple | Apple Leaf Black Spot | 81 | 1202 | 33.90% | Apple Leaf – Spot | 0.760 |
| Apple Leaf – Mosaic Virus | 224 | |||||
| Apple Leaf – Spot | 271 | |||||
| Apple Stem – Canker | 169 | |||||
| Apple Fruit – Black Rot | 65 | |||||
| Apple Leaf – Healthy | 215 |
| Plant Name | Class of Diseases | No. of Images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Apple | Scab Apple G | 241 | 1,769 | 6.61% | Healthy Apple | 0.871 |
| Scab Apple S | 174 | |||||
| Healthy Apple | 1354 | |||||
| Grape | Leaf-Blight Fungus Grape G | 70 | 3,167 | 11.80% | Leaf-Blight Fungus S | 0.903 |
| Black Rot Fungus Grape G | 435 | |||||
| Black Measles Fungus Grape G | 587 |
| Plant | Class of disease | No. of images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Rice | Bacterial Leaf Blight | 40 | 120 | 60% | Balanced | 0.00 |
| Brown Spot | 40 | |||||
| Leaf Blast | 40 | |||||
| Wheat | Leaf Rust | 40 | 80 | 40% | Balanced | 0.00 |
| Powdery Mildew | 40 |
| Plant | Class of disease | No. of images | Total Images (Crop) | % of Dataset | Dominant Disease Class | Skewness Measure |
| Rice | Bacterial Leaf Blight | 438 | 2710 | 73.05% | Brown Spot | 0.041 |
| Brown Spot | 466 | |||||
| Healthy | 464 | |||||
| Leaf Blast | 454 | |||||
| Leaf Scald | 448 | |||||
| Narrow Brown Spot | 440 |
| Dataset | Total Crops | Total Images | Avg. Crop-Level NIR | Inter-Crop Variance | Variance Level | Acquisition |
| PlantVillage | 14 | 54,306 | 0.77 | 0.97 | High | Lab |
| PlantDoc | 13 | 2,569 | 0.68 | 0.92 | High | Mixed |
| AI Challenger | 10 | 26,747 | 0.77 | 0.92 | High | Lab |
| NZDL Fruit Dataset | 5 | 3,545 | 0.53 | 0.63 | Moderate | Field |
| Turkey Plant Dataset | 7 | 4,447 | 0.79 | 0.95 | High | Field |
| FieldPlant Dataset | 3 | 8,629 | 0.99 | 0.18 | Moderate | Field |
| Imbalance Type | Impact on Model Learning | Potential Risk |
| Intra-Class Imbalance | Majority classes dominate gradient updates and decision boundaries | Low recall for minority diseases; biased classification |
| Inter-Crop Imbalance | Species with higher sample volume shape learned representations | Reduced cross-species generalization |
| Missing or Underrepresented Healthy Classes | Distorted separation between healthy and diseased states | Overconfident predictions in real-world deployment |
| Long-Tail Disease Distribution | Rare diseases insufficiently encoded in feature space | Poor detection of uncommon or emerging diseases |
| NIR Range | Structural Interpretation | Typical Model Behavior | Deployment Risk |
| 0.0 – 0.3 | Relatively balanced class distribution | Stable training; improved per-class recall; reduced need for heavy augmentation | Lower risk of minority-class suppression; better generalization potential |
| 0.3 – 0.7 | Moderate imbalance across disease categories | Bias toward majority classes; reliance on augmentation or class-weighting; possible overfitting | Reduced reliability for underrepresented diseases in field conditions |
| 0.7 – 1.0 | Severe skewness; dominance of few classes or absence of certain categories | Majority-class dominance; inflated overall accuracy; poor minority recall; architectural compensation (ensemble/hybrid models) | High risk of failure in real-field deployment, especially for rare but critical diseases |
| Ref | Dataset(s) Name | Learning Model | Data Augmentation Used | Model Accuracy (%) |
|---|---|---|---|---|
| [ 88 ] | PlantVillage Dataset | AlexNet | Yes | 86.5% |
| [ 90 ] | PlantVillage Dataset | Comparison between VGG16, MobileNetV2, InceptionV3, Xception and DenseNet121 (Modified Transfer Learning) | Yes | DenseNet121 achieved the highest accuracy (99.9%) |
| [ 91 ] | Proprietary Dataset (Mango Dataset collected from different fileds in Bangladesh) [ 92 ] | Comparison between VGG19, InceptionV3, ResNet152V2, DenseNet121, InceptionResNetV2, MobileNetV2 and Xception | Yes | InceptionV3 achieved the highest accuracy (99.87%) |
| [ 93 ] | PlantVillage Dataset | CNN Models - VGG16 and Xception | Yes | VGG with 98.18% |
| [ 10 ] | PlantVillage Dataset | AlexNet | NA | 95.75% |
| [ 74 ] | Proprietary Dataset (Images collected from Myanmar Grapevine Yard) | VGGNet 16 | Yes | 98.4 % |
| Ref | Dataset(s) Name | Learning Model | Data Augmentation Used | Model Accuracy (%) |
| [ 109 ] | PlantVillage Dataset | Ensemble of CNN (VGG-16, ResNet-50, DenseNet-121, Xception) and ViT (MobileViT, MViTv2, DeiT3, MaxViT) | Yes | Ensemble of DeiT3 and MaxViT achieved accuracy of 99.83% |
| [ 110 ] | Kaggle Wheat Disease Datasets | DenseNet201, MobileNetV2, Custom CNN (Ensemble - Bagging based Multi stage CNN) | No | 99.16% |
| [ 111 ] | PlantVillage Dataset | VGG16, VGG19, InceptionV3, ResNet101V2 (Ensemble Voting) | Yes | 92.3% |
| [ 107 ] | PlantVillage Dataset | United Ensemble Model - GoogLeNet and ResNet | Yes | 98.57% |
| [ 112 ] | PlantVillage Dataset | EfcientNetB0 and DenseNet 121 in ensemble mode. | No | 98.56% |
| [ 42 ] | Proprietary Dataset (The Turkey-Plant Dataset – images collected from the agricultural fields of the Bingol and Inonu Universities in Turkey) | AlexNet, GoogleNet, ResNet18, ResNet50, ResNet10 and DenseNet201. | No | 97.56% |
| Ref | Dataset(s) Name | Learning Model | Data Augmentation Used | Model Accuracy (%) |
|---|---|---|---|---|
| [ 114 ] | PlantVillage Dataset (Apple, Corn) | VGG16 + Inception-V3 + DenseNet201 + ViT (Hybrid (CNN Ensemble + Vision Transformer)) | Yes | 99.24% (Apple), 98% (Corn) |
| [ 72 ] | Proprietary Dataset (Images collected from fields) | Four CNN models VGG16, SSD, Faster R-CNN, and Yolov4 were used directly, with Yolov4 employed in a hybrid approach. | Yes | 96.47% |
| [ 83 ] | Proprietary Dataset (Images captured using multispectral camera) | CNN (ResNet) with Vision Transformers | Yes | 88.88% |
| [ 115 ] | Proprietary Dataset (Coffee [ 103 ] ) | SUNet (SegNet encoder + U-Net decoder + VGG16 backbone + Mask R-CNN + PSPNet) | Yes | 98.45% |
| [ 116 ] | PlantVillage Dataset | AlexNet and GoogleNet | NA | 97.62 % |
| [ 104 ] | PlantVillage Dataset | VGGNet 19 with EfficientNetB0, MobileNetV2 | Yes | 96.86 % |
| Ref | Dataset Name | Learning Model | Data Augmentation Used | Model Accuracy (%) |
| [ 76 ] | Proprietary Dataset (Images collected from a vineyard scanned with lab-scale HSI system in 900–1700 nm spectral range) | Partial least squares discriminant analysis (PLS-DA) | No | 97.17% |
| [ 124 ] | Combined Dataset (PlantVillage, VegNet, Bangladeshi Crop Disease, Eggplant Dataset [ 125 ] ) | PlantCareNet (5 convolution layers, 3 pooling layers, followed by fully connected layers before the final Softmax classifier.) | Yes | 97% |
| [ 122 ] | PlantVillage Dataset | Custom CNN with 3 convolution and 3 max-pooling layers followed by 2 fully connected layer | Yes | 91.2% |
| [ 123 ] | PlantVillage Dataset | Custom CNN incorporates several convolutional layers, with a pooling, activation, and fully connected layer. | No | 87.47% to 99.25% |
| [ 126 ] | PlantVillage Dataset | Custom lightweight 6-layer CNN | No | 98.4% |
| [ 64 ] | Proprietary Dataset (Images collected from a farmland) | Custom CNN model designed with two shallow convolutional layers followed by 3 Fully Connected layers, along with batch normalization and max-pooling layers. | Yes | 93.25% |
| Disease Type | Affected Plant Part/Appearance | Favorable Environmental Conditions |
| Apple Scab | Leaf and Fruit | Temperature on average: 20°C |
| Apple Powdery Mildew | Leaf and Fruit | Temperature ranging from 10°-25° C |
| Apple Marssonina Leaf Blotch | Leaf | High rainfall and temperatures between 20° and 22° C |
| Apple Alternaria Leaf Spot | Leaf | Temperature ranging from 25°-30° C |
| Apple Black Rot | Leaf and Fruit | 20°-22° C Temperature range and Moisture |
| Apple Sooty Blotch | Fruit | Temperature range 18°-27° C |
| Disease Type | Affected Plant Part/Appearance | Favorable Environmental Conditions |
| Apple Powdery Mildew | Leaf and Fruit | Humidity greater than 70% |
| Apple Collar Rot | Root | Over-watering and Moisture |
| Apple Seedling Blight | Root | High Humidity |
| Apple Core Rot | Fruit | High humidity and warm temperature |
| Apple Brown Rot | Fruit | Wet humid conditions |
| Disease Type | Affected Plant Part/Appearance | Favorable Environmental Conditions |
| Apple Scab | Leaf | Light infection: 9 hours at 18°C |
| Apple Alternaria Leaf Spot | Leaf | More than 6 hours of LWD |
| Challenge | Implication for Model Development and Deployment |
| Lack of synchronized environmental metadata | Prevents context-aware learning and limits multimodal benchmarking consistency. |
| Sensor heterogeneity across studies | Reduces reproducibility and complicates cross-dataset comparison. |
| Environmental variability across regions | Introduces domain shift and weakens model generalization in field deployment. |
| Absence of standardized multimodal benchmarks | Hinders fair architectural comparison and systematic performance evaluation. |
| Limited public availability of multimodal datasets | Restricts collaborative research and slows development of field-ready systems. |
| Open Challenge | Detailed Description and Implications |
|---|---|
| Scarcity of Large-Scale Real-Field Datasets | Although numerous plant disease datasets have been introduced, a substantial proportion remains laboratory-generated, characterized by uniform backgrounds, controlled illumination, and limited environmental noise. Real-field datasets are often smaller in size, geographically constrained, crop-specific, and frequently proprietary. This structural dependency on controlled datasets inflates experimental performance while limiting ecological validity and real-world generalization. The development of large-scale, geographically diverse, publicly accessible field datasets remains a pressing research priority. |
| Multi-Level Class Imbalance and Data Skewness | Class imbalance persists at multiple levels, including intra-class imbalance within disease categories, inter-crop imbalance across plant species, and cross-dataset imbalance due to inconsistent dataset design. Severe skewness biases model learning toward dominant classes, suppresses minority disease detection, and produces misleading overall accuracy metrics. While augmentation and resampling strategies are commonly applied, they do not fundamentally resolve structural imbalance inherent in dataset construction. More principled imbalance-aware learning and evaluation strategies remain underexplored in agricultural AI research. |
| Lack of Standardized Benchmarking Protocols | The absence of unified benchmarking frameworks complicates fair comparison across studies. Variability in train–test splits, preprocessing pipelines, augmentation policies, and performance metrics limits reproducibility. Additionally, architectural diversity across single, ensemble, hybrid, and custom models further hinders standardized evaluation. Without common validation protocols and cross-dataset testing strategies, reported improvements may reflect experimental configuration rather than genuine methodological advancement. |
| Limited Cross-Dataset Generalization Studies | Most existing research evaluates models on the same dataset used for training. Cross-dataset validation—training on one dataset and testing on another is rarely performed. This limits understanding of domain shift effects, acquisition variability, crop morphological differences, and environmental transferability. Robust real-world deployment requires models that generalize across datasets, regions, and acquisition conditions, yet systematic cross-dataset evaluation remains largely unexplored. |
| Environmental Context Integration Constraints | Environmental factors such as temperature, humidity, leaf wetness, and soil conditions significantly influence disease progression. However, most detection frameworks rely exclusively on visual inputs. Multimodal datasets combining image data with synchronized environmental measurements remain scarce. Key challenges include temporal alignment between sensor data and image capture, sensor calibration variability, missing environmental readings, and increased fusion complexity. Establishing standardized multimodal benchmarks is critical for advancing environmentally aware disease prediction systems. |
| Annotation Quality and Label Noise | Accurate annotation of plant disease images requires domain expertise. Field images may contain mixed symptoms, early-stage infections with subtle visual cues, or co-existing diseases on the same leaf. Inconsistent labeling protocols introduce annotation noise that negatively impacts model learning, generalization, and reproducibility. Establishing standardized annotation guidelines and expert-verified labeling pipelines is essential for improving dataset reliability. |