Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence
Organizations: Khalifa University Center for Autonomous Robotic Systems (KU-CARS), Khalifa University of Science and Technology, Abu Dhabi, United Arab Emirates. · Advanced Research and Innovation Center (ARIC), Khalifa University of Science and Technology, Abu Dhabi, United Arab Emirates. · Department of Computer Science and Engineering, Chung-Ang University, Seoul, South Korea. · Department of Aerospace Engineering, Khalifa University of Science and Technology, Abu Dhabi, United Arab Emirates.
Abstract
This review examines the evolution of plant phenotyping from conventional manual trait measurement to high-throughput, robotic, and artificial intelligence-driven crop monitoring. Despite significant advances in imaging, autonomous platforms, multimodal sensing, and deep learning, current phenotyping systems remain fragmented across sensing modalities, crop traits, growth stages, environments, and management objectives. We therefore frame phenotyping as an integrated \emph{seed-soil-plant-environment-management} (SSPEM) intelligence problem, where crop performance reflects interactions among seed quality, root-zone conditions, plant development, environmental exposure, and management actions. The review synthesizes conventional, high-throughput, robotic, and AI-driven phenotyping approaches, highlighting their capabilities and persistent limitations in temporal integration, multimodal reasoning, biological interpretation, and actionable decision support. Building on this analysis, we introduce a conceptual PhenoAgent framework that extends phenotyping beyond the estimation of isolated traits to evidence-based crop-state interpretation, uncertainty-aware reasoning, and management-oriented support. The PhenoAgent concept primarily brings together scattered advances in phenotyping to deliver insights ranging from detailed to high-level, such as what is happening in the crop, why it might be occurring, what evidence is missing, and what actions or additional measurements should be considered. We also discuss challenges in dataset scarcity, annotation, benchmarking, model generalization, and explainability. By linking multimodal phenotyping with agentic AI and closed-loop decision support, this review outlines a path to interpretable, scalable, and deployment-oriented crop intelligence.
Figures & tables
| Review | Focus | PP | SP | RZ | Rob | DL | FM | DT | AR | CLA | Limitations |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Pieruschka and Schurr [ 73 ] | History, infrastructure, standardization, and future needs in plant phenotyping | ✓ | ⚫ | ⚫ | ⚫ | ⚫ | ✗ | ✗ | ✗ | ✗ | Does not focus on agentic AI, foundation models, or closed loop seed soil plant reasoning. |
| Murphy et al. [ 64 ] | Deep learning in image based plant phenotyping | ✓ | ✗ | ✗ | ⚫ | ✓ | ⚫ | ✗ | ✗ | ✗ | Strong AI review, but primarily image based and not organized around multimodal decision intelligence. |
| Jin et al. [ 37 ] | Deep learning for high throughput seed phenotyping | ⚫ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | Seed specific focus, with limited connection to downstream plant, environment, management, and intervention decisions. |
| Wang et al. [ 102 ] | Image based high throughput phenotyping technologies and trends | ✓ | ⚫ | ⚫ | ✓ | ✓ | ⚫ | ⚫ | ✗ | ⚫ | Broad technology synthesis, but agentic reasoning and unified management intelligence are not central. |
| Zhang et al. [ 111 ] | Ground mobile robots for high throughput plant phenotyping from perception decision action perspective | ✓ | ✗ | ⚫ | ✓ | ✓ | ⚫ | ✗ | ⚫ | ✓ | Strong robotic pipeline focus, but less emphasis on seed soil plant digital twins and foundation agent reasoning. |
| Li et al. [ 49 ] | Foundation models in smart agriculture | ⚫ | ✗ | ⚫ | ⚫ | ✓ | ✓ | ⚫ | ⚫ | ⚫ | Broad smart agriculture focus, not specifically centered on plant phenotyping workflows and phenotypic validation. |
| RQ | Research question | Review contribution |
|---|---|---|
| RQ1 | How has plant phenotyping evolved from manual measurement to AI enabled automation? | Provides a compact historical synthesis of conventional phenotyping, high throughput sensing, robotic platforms, and AI driven trait extraction. |
| RQ2 | Which current technologies enable automated phenotyping? | Compares imaging, IoT, UAV, UGV, robotic, deep learning, foundation model, multimodal fusion, digital twin, and dashboard advances. |
| RQ3 | Where are current systems still fragmented? | Identifies gaps in dataset standardization, multimodal fusion, temporal reasoning, biological interpretability, and seed soil plant integration. |
| RQ4 | How can PhenoAgent define the next generation of phenotyping? | Introduces a conceptual framework for agentic (SSPEM) intelligence with explanation, forecasting, and intervention recommendation. |
| RQ5 | What is needed for high TRL translation? | Discusses benchmarking, validation, edge/cloud deployment, human oversight, data ownership, commercialization, and operational adoption. |
| Phenotyping layer | Typical conventional measurements | Common methods | Strengths | Key limitations | References |
|---|---|---|---|---|---|
| Seed quality and vigor | Viability, germination percentage, germination speed, seedling vigor, aging tolerance, stress germination | Standard germination test, accelerated aging, electrical conductivity, tetrazolium staining, seedling length and dry weight | Directly linked to crop establishment and seed lot value | Time consuming, often destructive or semi destructive, limited single seed tracking, weak field prediction under stress | Rajjou et al. [ 80 ] ; Marcos Filho [ 57 ] ; Reed et al. [ 82 ] ; Xing et al. [ 107 ] |
| Soil/substrate and root zone | Moisture, pH, electrical conductivity, salinity, nutrient availability, bulk density, organic matter, microbial or rhizosphere indicators | Soil sampling, laboratory chemical analysis, pot/substrate assays, root washing, soil cores, trenching | Provides mechanistic context for plant performance | Spatially sparse, delayed lab turnaround, destructive sampling, poor linkage to individual plant time series | Viscarra Rossel et al. [ 99 ] ; Bünemann et al. [ 10 ] ; de la Fuente Cantó et al. [ 18 ] ; Blanchy et al. [ 7 ] |
| Root architecture | Root depth, root angle, root length density, root biomass, crown root number, lateral roots, root hair traits | Excavation, shovelomics, root washing, rhizoboxes, minirhizotrons, root imaging after harvest | Captures belowground traits relevant to water and nutrient uptake | Laborious, genotype by soil sensitivity, disturbance of rhizosphere, low throughput in field plots | Trachsel et al. [ 94 ] ; Kuijken et al. [ 45 ] ; Atkinson et al. [ 6 ] ; Takahashi and Pradal [ 88 ] |
| Shoot growth and canopy traits | Plant height, leaf area, biomass, tillering, canopy cover, chlorophyll, flowering time, plant architecture | Rulers, calipers, SPAD meters, visual scoring, leaf area meters, destructive biomass harvest | Simple, interpretable, widely used in breeding and agronomy | Observer bias, sparse temporal resolution, destructive endpoints, limited spatial coverage | Fiorani and Schurr [ 26 ] ; Araus and Cairns [ 2 ] ; Watt et al. [ 103 ] ; Wang et al. [ 102 ] |
| Stress and disease response | Wilting, chlorosis, necrosis, disease severity, drought response, salinity response, nutrient deficiency symptoms | Visual scales, manual scoring, tissue sampling, physiological assays, yield loss assessment | Agronomically meaningful and accepted by breeders | Subjective scoring, late symptom visibility, low repeatability among observers, limited early warning ability | Poorter et al. [ 74 ] ; Tardieu et al. [ 89 ] ; Murphy et al. [ 64 ] |
| Yield and product quality | Grain yield, fruit number, fruit weight, harvest index, quality grades, postharvest traits | Manual harvest, combine yield, laboratory quality testing, grading | Final economic endpoint and breeding target | End of season only, confounded by many earlier processes, weak diagnostic value for cause of failure | Ray et al. [ 81 ] ; Cooper et al. [ 15 ] ; Mahmood et al. [ 56 ] |
| Year | Study | Domain | Contribution | Conventional relevance | Limitations | Relevance to PhenoAgent |
|---|---|---|---|---|---|---|
| 2011 | Furbank and Tester [ 28 ] | Plant phenomics | Framed phenotyping as a bottleneck limiting crop improvement | Established need for higher throughput measurement | Did not yet address modern agentic or multimodal reasoning | Provides historical motivation for automated intelligence |
| 2011 | Trachsel et al. [ 94 ] | Root phenotyping | Introduced shovelomics for field root architecture assessment | Practical method for root crown traits in field breeding | Destructive, labor intensive, limited temporal monitoring | Shows why belowground sensing must become less destructive |
| 2012 | Rajjou et al. [ 80 ] | Seed germination/vigor | Reviewed molecular and physiological bases of germination vigor | Connects seed vigor to establishment success | Molecular markers are not automatically linked to later crop management | Supports seed to plant phenotype continuity |
| 2013 | Fiorani and Schurr [ 26 ] | Plant phenotyping | Discussed future scenarios for noninvasive phenotyping | Emphasized resource use traits and environmental response | Large scale data integration remained a challenge | Anticipates the need for integrated phenotyping systems |
| 2014 | Araus and Cairns [ 2 ] | Field phenotyping | Positioned field HTPP as a crop breeding frontier | Linked phenotyping to breeding under realistic environments | Field data remain noisy and difficult to interpret causally | Motivates context aware phenotyping agents |
| 2015 | Marcos Filho [ 57 ] | Seed vigor | Reviewed seed vigor testing development and practice | Clarifies why vigor is more informative than germination alone | Many tests remain species dependent and laborious | Supports seed quality reasoning within PhenoAgent |
| Limitation | How it appears in conventional phenotyping | Scientific consequence | Operational consequence | PhenoAgent opportunity |
|---|---|---|---|---|
| Labor intensity | Manual scoring, plant measurement, seedling counting, root washing, destructive harvest, soil sampling | Limits population size, trait diversity, and temporal frequency | Increases cost and reduces commercial scalability | Automate observation using sensors, robots, and AI assisted annotation |
| Destructive endpoints | Biomass harvest, root excavation, tissue analysis, seedling dissection | Prevents continuous tracking of the same plant or plot | Delays understanding of stress progression and recovery | Link non-destructive sensing with selective ground-truth sampling |
| Subjectivity | Visual vigor, disease severity, wilting, lodging, chlorosis and canopy scores | Reduces reproducibility and introduces observer bias | Limits trust and comparability across sites and operators | Use calibrated models with uncertainty estimates and expert review |
| Sparse temporal data | Measurements collected weekly, at key stages, or only at harvest | Misses transient stress, early divergence, and treatment-response timing | Prevents timely intervention in greenhouse and field trials | Build temporal digital plant records and growth trajectories |
| Separated data streams | Seed, soil, weather, irrigation, image, and yield data stored independently | Weak causal interpretation across the seed soil plant continuum | Makes trial decisions reactive rather than predictive | Fuse seed, soil/root zone, plant, environment, and management data |
| Weak G E M coverage | Limited factorial combinations and few environments | Important interactions remain hidden or confounded | Slows breeding and management optimization | Use agentic reasoning to prioritize measurements, treatments, and interventions |
| Modality | Traits / Signals | Deployment | Strength | Current limitation | References |
|---|---|---|---|---|---|
| RGB imaging | Emergence, leaf area, canopy cover, color, organ count, lesions, fruit count, growth rate | Fixed cameras, UGVs, UAVs, smartphones, greenhouse stations | Low cost, high spatial detail, compatible with deep learning | Sensitive to illumination, occlusion, background, and visible-symptom delay | Minervini et al. [ 60 ] ; Pound et al. [ 75 ] ; Tausen et al. [ 90 ] ; Meraj et al. [ 59 ] |
| Thermal imaging | Canopy temperature, water stress, transpiration proxy, irrigation response | UAVs, fixed greenhouse cameras, gantry systems | Enables early water-status monitoring beyond RGB appearance | Requires calibration and correction for weather, angle, and emissivity | Perich et al. [ 72 ] ; Xie and Li [ 108 ] ; Nguyen et al. [ 65 ] |
| Multispectral imaging | Vegetation indices, chlorophyll, biomass proxy, nitrogen, stress indices | UAVs, handheld sensors, greenhouse cameras | Balances cost and physiological sensitivity | Band selection and vegetation-index saturation can limit generality | Sankaran et al. [ 84 ] ; Xie and Li [ 108 ] ; Khuimphukhieo and Silva [ 42 ] |
| Hyperspectral imaging | Pigment, water, nitrogen, disease, salinity and biochemical fingerprints | UAVs, proximal carts, laboratory and greenhouse imaging | Dense spectral information for subtle stress and composition changes | High data volume, calibration burden, limited interpretability without models | Nguyen et al. [ 65 ] ; de Silva and Brown [ 19 ] ; Wang et al. [ 102 ] |
| Depth and LiDAR | Plant height, canopy volume, row structure, lodging, 3D architecture, biomass proxy | UAV LiDAR, UGV LiDAR, gantry systems, RGB-D cameras | Adds spatial structure and geometry, reducing ambiguity of 2D images | Cost, registration, occlusion, point-cloud processing, and cross-platform calibration | Busemeyer et al. [ 11 ] ; Xie and Li [ 108 ] ; Nguyen et al. [ 65 ] |
| IoT and environmental sensors | Temperature, humidity, light, CO2, substrate moisture, pH, EC, irrigation, fertigation, actuator logs | Smart greenhouses, vertical farms, field sensor networks | Provides continuous causal context for plant response | Often stored separately from image-derived traits and biological labels | Wolfert et al. [ 105 ] ; Peladarinos et al. [ 71 ] ; Rahman et al. [ 79 ] |
| Year | Study | Domain | Main advance | Contribution | Remaining gap | Relevance to PhenoAgent |
|---|---|---|---|---|---|---|
| 2011 | Furbank and Tester [ 28 ] | Phenomics vision | Defined phenotyping bottleneck | Motivated sensor-based scale-up | Limited AI and robotics integration | Historical baseline for automation |
| 2012 | White et al. [ 104 ] | Field phenomics | Framed field-based phenomics for genetics | Linked field measurements to genetic discovery | Sparse automation compared with current systems | Supports field-relevant validation |
| 2013 | Busemeyer et al. [ 11 ] | Multi-sensor platform | Tractor-based BreedVision system | Combined sensors for field breeding | Heavy platform, limited reasoning layer | Demonstrates platform-level integration |
| 2013 | Cobb et al. [ 14 ] | Crop improvement | Defined next generation phenotyping requirements | Connected phenotyping to genotype-phenotype understanding | Data integration remained unresolved | Supports biological interpretation focus |
| 2014 | Andrade-Sanchez et al. [ 1 ] | Field robotics | Field-based high throughput phenotyping platform | Enabled repeated crop measurements in field plots | Platform-specific workflow | Motivates mobile phenotyping infrastructure |
| 2014 | Araus and Cairns [ 2 ] | Field HTPP | Positioned field phenotyping as breeding frontier | Linked sensing to selection under real environments | Translating data to knowledge remained hard | Supports breeding-facing PhenoAgent |
| Current Capability | What current systems do well | Main unresolved gap | Design requirement for PhenoAgent | Supporting literature |
|---|---|---|---|---|
| High throughput sensing | Rapid imaging, repeated measurements, spectral and structural traits | Sensor outputs remain fragmented by modality and platform | Time-synchronized sensor fusion by plant, plot, treatment, and growth stage | Tardieu et al. [ 89 ] ; Wang et al. [ 102 ] |
| Robotic and UAV platforms | Scalable field or greenhouse data acquisition | Limited biological interpretation and closed loop action | Platform-aware sensing plans and autonomous follow up measurements | Atefi et al. [ 5 ] ; Zhang et al. [ 111 ] ; Khuimphukhieo and Silva [ 42 ] |
| Deep learning perception | Detection, segmentation, classification, and regression | Task-specific models with limited transfer and weak causal explanation | Foundation-model perception with uncertainty and human correction | Pound et al. [ 75 ] ; Murphy et al. [ 64 ] ; Li et al. [ 49 ] |
| Multimodal fusion | Improved prediction by combining spectral, structural, thermal, and environmental data | Fusion often remains feature-level rather than biology-level | seed soil plant environment management representation learning | Nguyen et al. [ 65 ] ; Coppens et al. [ 16 ] ; Papoutsoglou et al. [ 69 ] |
| Digital twins and dashboards | Monitoring, simulation, prediction, and user visualization | Many implementations are descriptive dashboards or digital shadows | Agentic digital twin with reasoning, evidence tracking, and decision support | Purcell and Neubauer [ 77 ] ; Rahman et al. [ 79 ] ; Ojo et al. [ 66 ] |
| Decision support | Alerts, trends, forecasts, and rule-based recommendations | Weak explanation of causes and next-best measurements | Explainable recommendations grounded in crop physiology and trial context | Walter et al. [ 101 ] ; Cesco et al. [ 13 ] ; Arshad et al. [ 3 ] |
| Step | Stage | Main inputs | Core function | Expected output |
|---|---|---|---|---|
| 1 | Phenotyping context definition | Crop type, growth stage, target trait, treatment, management objective, and trial metadata | Define what should be monitored, why it matters, and how crop units will be tracked over time | Crop-unit identity, trial context, and phenotyping objective |
| 2 | Multimodal data acquisition | Fixed cameras, UAV/UAD platforms, UGV robots, IoT sensors, seed records, substrate/root-zone measurements, and management logs | Collect complementary visual, environmental, biological, and management evidence from the crop system | Raw multimodal crop evidence |
| 3 | Temporal alignment and quality control | Images, sensor logs, timestamps, crop-unit IDs, and metadata | Synchronize data streams, remove low-quality records, align observations with plant/tray/plot identity, and organize measurements along the crop timeline | Clean time-aligned evidence record |
| 4 | Trait and state extraction | RGB/depth/thermal/spectral images, IoT signals, substrate data, and management events | Estimate crop traits and state indicators such as emergence, vigor, canopy growth, stress symptoms, disease cues, root-zone condition, and yield-related proxies | Trait/state profile with confidence values |
| 5 | Plant digital twin construction | Trait profile, environmental history, root-zone status, seed information, and management records | Build a time-resolved crop-state representation that links phenotype, environment, substrate/root-zone condition, and management history | Dynamic plant/tray/plot digital twin |
| 6 | PhenoAgent reasoning | Digital twin, perception outputs, crop knowledge, retrieved evidence, forecasting models, uncertainty estimates, and expert feedback | Interpret the crop state, identify likely stress drivers, compare alternative explanations, and determine whether additional evidence is needed | Explainable crop-state interpretation |
| Dimension | Current Common Practice | PhenoAgent Direction | Expected Advantage | Key References |
|---|---|---|---|---|
| System goal | Estimate traits, indices, or labels from images and sensors | Explain crop state and recommend evidence-aware next steps | Moves from measurement to actionable intelligence | Tardieu et al. [ 89 ] ; Murphy et al. [ 64 ] |
| Data organization | Separate image folders, sensor logs, trial sheets, and dashboard values | Unified seed soil plant environment management time series | Enables causal interpretation across crop development | Coppens et al. [ 16 ] ; Papoutsoglou et al. [ 69 ] |
| AI model role | Task-specific detection, segmentation, classification, or regression | Modular perception plus foundation model and retrieval-assisted reasoning | Supports reusable and context-aware interpretation | Bommasani et al. [ 8 ] ; Li et al. [ 49 ] |
| Decision support | Threshold alerts, plots, and descriptive dashboards | Hypothesis ranking, uncertainty, intervention recommendation, and follow up sensing | Improves grower and breeder decision confidence | Walter et al. [ 101 ] ; Rahman et al. [ 79 ] |
| Human role | Manual labeling, scouting, interpretation, and final decision | Expert guided validation, feedback, correction, and approval of interventions | Preserves agronomic expertise while reducing routine burden | Wolfert et al. [ 105 ] ; Peladarinos et al. [ 71 ] |
| Translation path | Research prototype or sensor-specific system | Scalable proof of concept progressing toward high-TRL deployment | Converts data infrastructure into deployable phenotyping product | Atefi et al. [ 5 ] ; DigiHortiRobot [ 25 ] |
| Gap | Current challenge | Future opportunity | Validation requirement | Key references |
|---|---|---|---|---|
| Dataset scarcity | Few datasets cover multiple crops, sensors, stages, environments, and management histories | Longitudinal seed soil plant datasets with real and synthetic data | Cross-site benchmark splits and metadata completeness checks | Papoutsoglou et al. [ 69 ] ; Liu et al. [ 52 ] |
| Annotation burden | Expert labels are expensive and inconsistent across traits and crops | Promptable segmentation, active learning, weak labels, expert-in-the-loop correction | Label agreement, uncertainty reporting, and audit trails | Kirillov et al. [ 43 ] ; Murphy et al. [ 64 ] |
| Benchmarking | Many studies report isolated accuracy on private datasets | Open protocols for detection, trait extraction, stress prediction, and decision value | Standard metrics for accuracy, robustness, calibration, and agronomic utility | David et al. [ 17 ] ; Jiang and Li [ 36 ] |
| Data standards | Image, sensor, environment, and management records are often disconnected | FAIR and MIAPPE-aligned multimodal phenotyping records | Reusable schemas linked to plant/tray/plot identity and timestamps | Coppens et al. [ 16 ] ; Papoutsoglou et al. [ 69 ] |
| Biological reasoning | Models detect patterns but often do not explain crop causes | Agentic multimodal reasoning with crop knowledge and digital twins | Expert agreement, causal plausibility, and intervention outcomes | Tardieu et al. [ 89 ] ; Li et al. [ 49 ] |
| Barrier | Why it matters | Design response | TRL validation metric | Related references |
|---|---|---|---|---|
| Sensor robustness | Dust, humidity, lighting, drift, and occlusion reduce data quality | Calibration routines, QA flags, redundant sensing, automated health checks | Uptime, missing-data rate, calibration error | Xie and Li [ 108 ] ; Nguyen et al. [ 65 ] |
| Compute constraints | Real-time alerts and robotics need low latency, while training requires scale | Edge inference for alerts; cloud for training, dashboard, and historical analysis | Latency, throughput, energy use, cloud cost | Kalyani et al. [ 39 ] ; Rahman et al. [ 79 ] |
| Model maintenance | Crop cycles, sensors, and facilities change over time | Continuous validation, drift monitoring, versioned models, human correction | Degradation rate and retraining interval | Murphy et al. [ 64 ] ; Li et al. [ 49 ] |
| Data governance | Trial and yield records may be commercially sensitive | Access control, local/cloud policy, anonymization, audit logs, ownership agreement | Compliance, user acceptance, exportability | Wolfert et al. [ 105 ] ; Walter et al. [ 101 ] |
| Workflow integration | Systems fail if they interrupt grower or breeder routines | Dashboard co-design, role-specific views, mobile alerts, simple feedback tools | Usability, adoption, decision response time | Peladarinos et al. [ 71 ] ; Cesco et al. [ 13 ] |
| Commercial readiness | Research prototypes often lack support, reliability, and business model | TRL staging, pilot validation, service model, training, maintenance plan | TRL milestone completion and decision-value evidence | DigiHortiRobot [ 25 ] ; Zhang et al. [ 111 ] |