Organizations: College of Informatics, Huazhong Agricultural University, Wuhan 430070, China · Yunnan Modern Agricultural Industry Research Institute Co., Ltd.,Kunming, Yunnan 650000, China · College of Biological and Agricultural Engineering, Jilin University, Changchun 130022, China · Key Laboratory of Efficient Sowing and Harvesting Equipment, Ministry of Agriculture and Rural Affairs, Jilin University, Changchun 130022, China
Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.
Figures & tables
Figure 1: Framework of the proposed KAN-based safe reinforcement learning greenhouse climate control system.
Figure 2: Greenhouse-crop system dynamic relationship diagram.
State x(t)
Output y(t)
x1(t)
Dry-weight ( kg/m2 )
y1(t)
Dry-weight ( g/m2 )
x2(t)
Indoor CO2 ( kg/m3 )
y2(t)
Indoor CO2 (ppm)
x3(t)
Indoor temperature ( ∘ C)
y3(t)
Indoor temperature ( ∘ C)
x4(t)
Indoor humidity ( kg/m3 )
y4(t)
Indoor humidity (%)
Control input u(t)
Disturbance d(t)
u1(t)
CO2 injection ( mg/m2/s )
d1(t)
Radiation ( W/m2 )
Table 1: Definitions of state, input, output, and disturbance variables
Item
Value
Observation dimension nobs
62
KAN layers (actor)
3: [62,128,128,128]
KAN layers (critic)
3: [62,128,128,128]
Spline order k
3
Grid size
3
Grid range
[−1,1]
Table 2: Network and implementation settings of the KAN-based actor and critic networks.
Variable
Value
Unit
Description
cCO2
0.1906
€/kg
CO 2 price coefficient
cheat
0.1281
€/kWh
Heating price coefficient
cDW
22.29
€/kg
Crop dry-weight price
λCO2
5×10−5
–
Weight for CO 2 violations
λTmin
3×10−3
–
Weight for lower temperature violations
λTmax
5×10−3
–
Weight for upper temperature violations
Table 3: Economic parameters and constraint cost coefficients.
Variable
Lower
Upper
Unit
Climate variables
yCO2
500
1600
ppm
yT
10
20
∘ C
yRH
0
80
%
Control inputs
uCO2
0
1.2
mg/m 2 /s
Table 4: Climate variable ranges and control input limits.
Hyperparameter
Value
Total training timesteps
4×106
Number of parallel environments ( nenvs )
8
Rollout length ( nsteps )
1920
Batch size
1920
Optimization epochs per update
10
Discount factor ( γ )
0.98
Table 5: PPO hyperparameters used in the experiments.
Hyperparameter
Symbol
Value
Initial penalty coefficient
λ0
1.0
Minimum penalty coefficient
λmin
0.05
Maximum penalty coefficient
λmax
50.0
Constraint threshold
d
0.7
Penalty update interval (episodes)
K
4
Penalty learning rate
αλ
0.01
Table 6: RCPO hyperparameters for constraint regulation.
Method
Cumulative profit
Δ Profit (%)
Cumulative violation
Δ Violation (%)
PPO (baseline)
3.782 ( ±0.146 )
– ( ±3.86 )
0.858 ( ±0.027 )
– ( ±3.15 )
RCPO w/o KAN & Time
3.543 ( ±0.061 )
−6.32 ( ±1.72 )
0.695 ( ±0.026 )
−19.00 ( ±3.74 )
RCPO w/o Time
3.635 ( ±0.064 )
−3.89 ( ±1.76 )
0.712 ( ±0.039 )
−17.02 ( ±5.48 )
RCPO w/o KAN
3.755 ( ±0.165 )
−0.71 ( ±4.39 )
0.710 ( ±0.025 )
−17.25 ( ±3.52 )
KAN-RCPO-PPO
3.892 ( ±0.147 )
+2.91 ( ±3.78 )
0.698 ( ±0.029 )
−18.65 ( ±4.15 )
Table 7: Performance comparison under different controller configurations.
Figure 3: Ablation study on policy representation enhancements. Training performance comparison between baseline PPO, PPO with cyclic time features, and PPO with KAN-based policy networks. All methods are trained without RCPO constraint regulation.
Figure 4: Normalized cumulative climate constraint violations for different climate variables achieved by the proposed RCPO-PPO controller relative to the baseline PPO controller. The baseline PPO result is normalized to 100% .
Figure 5: Normalized cumulative operation costs associated with heating and CO2 input achieved by the proposed RCPO-PPO controller relative to the baseline PPO controller. The baseline PPO result is normalized to 100% .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Parameter
Value
Parameter
Value
p1
0.544
p11
7.50×10−6
p21
3.60×10−3
p2
2.65×10−7
p12
8.31
p22
9348
p3
53
p13
273.15
p23
8314
p4
3.55×10−9
p14
101325
p24
273.15
p5
5.11×10−6
p15
0.044
p25
17.4
p6
2.30×10−4
p16
3.00×104
p26
239
Appendix
Table 8: Nominal parameter values of the lettuce greenhouse model.
Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g. for safety-critical domains) has not been carefully examined. In this work, we introduce Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization. We show that the standard practice of scalarizing rewards before normalization introduces a critical Lagrangian-specific failure mode: GRPO's within-group normalization makes constrained optimization highly sensitive to how multi-component learning signals are aggregated. We show that scalarizing rewards before normalization introduces shared-denominator coupling, so that changing one multiplier alters not only the emphasis on its corresponding constraint, but also the relative weighting of the reward and other constraints. We address this with a simple but crucial modification: scalarizing standardized advantages rather than rewards. This yields a better-conditioned update by addressing the coupling induced by reward scalarization, resulting in better-behaved multiplier dynamics and more stable constraint enforcement in practice. Empirically, across a controlled gridworld, a real-world autonomous driving benchmark, and a mathematical reasoning task, Constrained GRPO consistently achieves better adherence to specified constraints while maintaining or improving task performance.
Our preliminary experiments on gym-DSSAT maize irrigation tasks revealed that +/-2 degrees C temperature noise causes an 11.9% reduction in economic returns for PPO policies trained under clean conditions - a systematic robustness deficit that existing research has not adequately addressed. This paper tackles three interconnected limitations impeding practical deployment of agricultural RL systems: the trade-off between early-stage learning efficiency and late-stage generalization capability; the naive additive combination of intrinsic and extrinsic rewards in exploration-augmented PPO; and uniform measurement noise injection strategies that disregard empirically validated differential sensitivity across agricultural state variables. We introduce three systematic innovations: Progressive Generalization Augmentation (PGA) implementing a three-phase curriculum (clean training 0-800 episodes, progressive 800-1200, full augmentation 1200-2000); a deeply coupled RND-PPO architecture with dual-channel GAE normalization, progress-decayed intrinsic coefficients, and semantic discretization; and domain-prioritized noise injection with hierarchical activation. Our experimental evaluation demonstrates: 8.43% yield improvement and 16.42% nitrogen use efficiency improvement over SOTA BERT-DQN in Florida; 5.61% yield improvement in Zaragoza (though 3.67% lower economic score due to challenging Mediterranean climate); and 94.4% vs 80.0% performance retention under combined perturbations. All experiments used 5 random seeds on NVIDIA A100 GPUs with 4.2+/-0.3 hours per run (2000 episodes, 2048-step buffer, 64 mini-batch size).
The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision has begun, and the Lagrange multiplier of PPO-Lagrangian grows only after the episode budget has been exceeded. At race speeds, where collisions are instantaneous and irreversible, any safety mechanism that waits for cost to accumulate is structurally too late. We present VLM-Safe-RL, a framework that integrates a frozen vision-language model into the CMDP Lagrangian update as an anticipatory cost term. The framework comprises four contributions: (i) Decoupled Dual-Path CLIP, independent reward/cost paths that respect the CMDP's factorization; (ii) VLM-Lagrange, an augmented multiplier update that incorporates a per-step VLM cost as an anticipatory term; (iii) Confidence Gating, a Bayes-optimal weight derived from a logistic noise model on the CLIP margin; and (iv) VLMPPOLag, the composed algorithm. On Safety-Gymnasium FormulaOne L2, our principal evaluation (n=5 seeds, 106 steps, budget dlim=25) VLMPPOLag+Conf is the only configuration in our default budget comparison that simultaneously retains substantive return (Jr≈40) and holds cost within budget on a majority of seeds; the five constraint-aware baselines (PPOLag, CPO, CPPOPID, CPO-CLG, PPOLag-RND) each fail at least one requirement. The mechanism generalizes to held-out MetaDrive Medium (catastrophe rate 41%→26%, 95% bootstrap CI [−26,−5],pp) and shows directionally consistent transfer to Bullet Safety-Gym; we report honestly where it does not (MetaDrive Easy/Hard, Qwen2-VL backbone) and trace the Hard failure to a Lagrangian-regulation pathology rather than the VLM signal itself. To our knowledge, this is the first work to use frozen VLM signals as an anticipatory cost term inside the CMDP Lagrangian update.