Organizations: Graduate School of Agricultural and Life Sciences, The University of Tokyo, 1-1-1 Midori-cho, Nishitokyo, Tokyo 188-0002, Japan · Engineering Research Center of Plant Phenotyping, Ministry of Education; Jiangsu Collaborative Innovation Center for Modern Crop Production; Academy for Advanced Interdisciplinary Studies, Nanjing Agricultural University, Nanjing 210095, China · National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China · Institute of Agricultural Machinery, NARO, 3-1-3 Kannondai, Tsukuba, Ibaraki 305-8604, Japan · Institute of Life and Environmental Sciences, University of Tsukuba, 1-1-1 Tennodai, Tsukuba, Ibaraki 305-8572, Japan · Next Generation Artificial Intelligence Research Center, The University of Tokyo, Tokyo, Japan
Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, organs, growth stages, and field conditions. General-purpose vision foundation models offer a natural route to label efficiency, but their web-scale pretraining objectives transfer weakly to agricultural imagery, where semantics are often determined by fine organ geometry inside repetitive, texture-dominated scenes. We introduce SPROUT, a diffusion foundation model for multi-crop plant phenotyping. SPROUT learns from 2.6 million unlabeled open-field images (MCD-2.6M) using a pixel-space Diffusion Transformer, and selects transferable features with a label-free effective-rank criterion over denoising timesteps. This design shifts pretraining from crop-based invariance to structure-preserving denoising, making the representation better aligned with dense phenotyping tasks. We evaluate SPROUT across dense phenotyping tasks, including organ segmentation, crop-weed parsing, depth estimation, and counting. SPROUT consistently improves over strong web-pretrained baselines, with the largest gains on dense structural prediction, and shows favorable label and compute efficiency compared with general-purpose and crop-specific foundation models. The source code and MCD-2.6M dataset are publicly available.
Figures & tables
Figure 1 : Left: PCA visualization of SPROUT feature maps. In the SPROUT embedding space, different plant organs exhibit clear semantic separation, while the same organ shares consistent semantics. SPROUT captures and understands the structural information of plants. Right: SPROUT’s performance on agricultural vision tasks. Absolute Relative Error is the metric for depth estimation, Mean Square Error is the metric for counting, and Intersection over Union (IoU) is used for all other tasks. Metrics where lower values indicate better performance are inversely normalized. SPROUT significantly outperforms general-purpose VFMs across a wide range of tasks, particularly in structural understanding and dense prediction tasks.
Figure 2 : Comparison of dense features. We employ Principal Component Analysis (PCA) to reduce the dimensionality of the feature maps to three dimensions, projecting them into RGB space ( Rh×w×c→Rh×w×3 ). Size of all feature maps is 128×128 . Compared to prior methods, SPROUT yields clearer features with less noise and distinct semantics.
Method
Pretraining dataset
Model
Apple
Peach [ 37 ]
Pear [ 37 ]
Grape [ 33 ]
Wheat [ 42 ]
Rice [ 46 ]
Flower [ 37 ]
Fruit [ 13 ]
Flower
Flower
Fruit
Spike
Stem
Leaf
Green Veg
Senescent Veg
Panicle
Self-Distillation
DINOv2 [ 25 ]
LVD-142M
ViT-S-16
55.71
67.90
57.86
63.12
88.48
80.25
32.22
80.71
80.71
47.54
72.77
ViT-B-16
58.44
69.18
56.96
63.23
88.85
81.05
34.80
80.96
80.74
47.77
73.67
ViT-L-16
54.91
69.43
55.16
61.70
88.72
81.77
37.05
81.46
81.19
48.74
75.04
DINOv3 [ 35 ]
LVD-1689M
ViT-S-16
45.71
64.97
58.18
59.43
87.64
76.47
25.17
77.09
78.88
46.84
71.73
Table 1 : Organ-level semantic segmentation Intersection-over-Union (IoU) across multiple crops and organ categories. SPROUT consistently achieves the highest IoU across all crops and organs.
Method
Bean
Carrot
Maize
Pea
Potato
Pumpkin
Rice
Soybean
SugarBeet
Sunflower
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
Crop
Weed
DINOv2-L
79.98
63.50
82.97
71.07
78.91
41.59
61.15
29.71
86.33
29.47
99.43
47.60
79.56
36.34
67.71
25.15
78.32
52.52
86.49
51.28
DINOv3-L
78.31
62.28
83.00
70.74
77.10
36.22
60.00
23.92
85.53
30.31
84.33
47.63
79.23
40.30
64.27
19.36
79.03
44.90
85.97
51.36
MSN-L
74.09
55.92
75.94
60.04
72.00
29.32
54.77
14.60
81.45
23.60
80.32
31.82
77.17
30.01
58.44
9.90
71.67
23.04
78.14
41.57
CLIP-L
80.96
64.67
84.20
73.02
79.88
42.22
61.53
30.72
87.41
34.46
86.63
51.03
80.70
39.54
68.33
28.87
80.96
51.84
86.65
51.35
SigLIP-L
77.49
59.23
81.46
68.93
74.43
35.71
58.23
19.90
84.70
25.80
81.82
40.15
77.87
36.68
62.75
20.81
77.75
43.08
84.46
49.09
Table 2 : Plant-level segmentation results on ten crop datasets. The table reports class-wise IoU for crop and weed segmentation. SPROUT achieves leading performance across most categories, indicating broad cross-species transfer.
Method
Pretraining dataset
Model
Params
Pretraining cost ↓
Stem IoU ↑
mIoU ↑
FOMO4Wheat
ImAg4Wheat-2.5M
ViT-G-16
1100 M
9216 A100 Hours
46.85
74.63
SPROUT
MCD-2.6M
UDiT-S
51 M
245 A100 Hours
48.84
74.77
UDiT-B
112 M
525 A100 Hours
52.88
76.46
UDiT-L
361 M
1440 A100 Hours
58.55
78.38
Table 3 : Comparison with the wheat foundational model FOMO4Wheat on wheat organ segmentation. One A100 hour refers to one hour of compute time using a single NVIDIA A100 GPU. Our SPROUT models obtain higher performance with markedly lower pretraining cost.
Method
AbsRel ↓
MAE ↓
RMSE ↓
DINOv2-L
0.0060
0.0066
0.0103
DINOv3-L
0.0062
0.0068
0.0110
CLIP-L
0.0074
0.0078
0.0139
SPROUT-L
0.0045
0.0050
0.0071
Table 4: Quantitative comparison of depth estimation performance on the Sugar Beet dataset.
Figure 7
Figure 5 : Scaling the pretraining dataset size. Left, downstream performance steadily improves as the unlabeled dataset scales up to 6.4×104 samples. After that, performance saturates, and adding more homogeneous data yields diminishing returns. Middle, model convergence under different dataset sizes. Right, convergence analysis of SPROUT-L across dataset scales. The required training steps scale linearly with the square root of dataset size, enabling estimation of the optimal number of training iterations.
Semantic segmentation in agricultural imagery is often evaluated under in-domain protocols, yet practical deployment requires robustness to appearance perturbations, limited annotations, and cross domain shift. This paper presents a diffusion-guided hybrid segmentation framework in which U-Net, DeepLabV3+, and SegFormer backbones generate coarse masks that are refined by Denoising Diffusion Probabilistic Models (DDPM), latent diffusion, or semantic-guided diffusion. The framework is evaluated through a 3x3 architectural screening study on PlantSegV3, followed by boundary-constrained optimization, perturbation-guided retraining, low-data evaluation, constrained hyperparameter screening, and controlled cross-domain adaptation. On PlantSegV3, the best selected hybrid model achieves 71.83% refined mean Intersection-over-Union (mIoU) and 26.10% refined Boundary-F1, and the selected models remain stable under substantially reduced supervision, demonstrating strong annotation efficiency. Perturbation analysis identifies grayscale conversion, fog, coarse dropout, and shadow as the most disruptive appearance shifts, and the resulting augmentation policy substantially improves robustness during retraining. The adapted models further show effective transfer to external agricultural datasets under limited target supervision, indicating that diffusion refinement and boundary-aware optimization provide transferable structural priors. Overall, the results show that carefully matched backbone-refiner pairings, combined with perturbation-aware retraining, can improve structural delineation and robustness under realistic resource and distribution constraints.
Gurbhit Chaurakoti, Soumyashree Kar
Department of Electrical Engineering, National Institute of Technology Delhi, New Delhi, 110036, India · Centre of Studies in Resources Engineering, Indian Institute of Technology Bombay, Powai, Mumbai, 400076, India
3D plant phenotyping is notoriously known to be procedure-complicated and of low throughput due to the extensive multi-view imaging, the fragile 3D reconstruction pipeline, and the additional cost from reconstructed geometry to phenotypic extraction. These limitations are further amplified in low-cost data acquisition, where smartphone videos or sparsely sampled multi-view images provide limited view overlap and self-occlusion. In this work, we show that the conventional 3D plant phenotyping pipeline could be streamlined and significantly accelerated with 3D Foundation Models (3DFMs), and particularly, present one of the first cross-crop 3D phenotyping frameworks powered by 3DFMs. The framework replaces COLMAP-style sparse initialization with 3DFM-based feed-forward geometric recovery, combines geometry-constrained 3D Gaussian Splatting for dense reconstruction, enables few-view reconstruction through iterative view synthesis and refinement, and converts reconstructed geometry into measurable organs through 2D-to-3D semantic transfer, metric scale recovery, and organ instance separation. We further construct a cross-crop dataset with smartphone-based image acquisition, diverse plant morphologies, and manual annotations for segmentation and phenotypic evaluation. Experiments across 26 plant sequences show that 3D Foundation Models reduce the average reconstruction time from 6.52 minutes to 1.58 seconds while maintaining high reconstruction quality and phenotyping accuracy. These results suggest a fresh technical route for high-throughput 3D plant phenotyping, from low-cost image acquisition to fast reconstruction, perception, scale recovery, and phenotypic measurement.
Hanyue Jia, Wei Zhou, Wenbo Zhou +3
Northwest A&F University, Yangling, 712100, China · Huazhong University of Science and Technology, Wuhan, 430074, China · Wuhan Institute of Technology, Wuhan, 430205, China
High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck," where manual data collection is labor-intensive and prone to observer bias. Conventional closed-set computer vision systems fail to address this challenge, as they require extensive species-specific annotation and lack the flexibility to handle diverse breeding populations. To bridge this gap, we present CropVLM, a Vision-Language Model (VLM) adapted for the agricultural domain via Domain-Specific Semantic Alignment (DSSA). Trained on 52,987 manually selected image-caption pairs covering 37 species in natural field conditions, CropVLM effectively maps agronomic terminology to fine-grained visual features. We further introduce the Hybrid Open-Set Localization Network (HOS-Net), an architecture that integrates CropVLM to enable the detection of novel crops solely from natural language descriptions without retraining. By eliminating the reliance on species-specific training data, CropVLM provides a scalable solution for high-throughput phenotyping, accelerating genetic gain and facilitating large-scale biodiversity research essential for sustainable agriculture. The trained model weights and complete pipeline implementation are publicly available at: https://github.com/boudiafA/CropVLM. In comprehensive evaluations, CropVLM achieves 72.51% zero-shot classification accuracy, outperforming seven CLIP-style baselines. Our detection pipeline demonstrates superior zero-shot generalization to novel species, achieving 49.17 AP50 on our CVTCropDet benchmark and 50.73 AP50 on tropical fruit species, compared to 34.89 and 48.58 for the next-best method, respectively.
Abderrahmene Boudiaf, Sajd Javed
Department of Electrical Engineering and Computer Science, Khalifa University of Science and Technology, Abu Dhabi, United Arab Emirates