Organizations: College of Informatics, Huazhong Agricultural University, Wuhan 430070, China · Yunnan Modern Agricultural Industry Research Institute Co., Ltd.,Kunming, Yunnan 650000, China · College of Biological and Agricultural Engineering, Jilin University, Changchun 130022, China · Key Laboratory of Efficient Sowing and Harvesting Equipment, Ministry of Agriculture and Rural Affairs, Jilin University, Changchun 130022, China
Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.
Figures & tables
Figure 1: Framework of the proposed KAN-based safe reinforcement learning greenhouse climate control system.
Figure 2: Greenhouse-crop system dynamic relationship diagram.
State x(t)
Output y(t)
x1(t)
Dry-weight ( kg/m2 )
y1(t)
Dry-weight ( g/m2 )
x2(t)
Indoor CO2 ( kg/m3 )
y2(t)
Indoor CO2 (ppm)
x3(t)
Indoor temperature ( ∘ C)
y3(t)
Indoor temperature ( ∘ C)
x4(t)
Indoor humidity ( kg/m3 )
y4(t)
Indoor humidity (%)
Control input u(t)
Disturbance d(t)
u1(t)
CO2 injection ( mg/m2/s )
d1(t)
Radiation ( W/m2 )
Table 1: Definitions of state, input, output, and disturbance variables
Item
Value
Observation dimension nobs
62
KAN layers (actor)
3: [62,128,128,128]
KAN layers (critic)
3: [62,128,128,128]
Spline order k
3
Grid size
3
Grid range
[−1,1]
Table 2: Network and implementation settings of the KAN-based actor and critic networks.
Variable
Value
Unit
Description
cCO2
0.1906
€/kg
CO 2 price coefficient
cheat
0.1281
€/kWh
Heating price coefficient
cDW
22.29
€/kg
Crop dry-weight price
λCO2
5×10−5
–
Weight for CO 2 violations
λTmin
3×10−3
–
Weight for lower temperature violations
λTmax
5×10−3
–
Weight for upper temperature violations
Table 3: Economic parameters and constraint cost coefficients.
Variable
Lower
Upper
Unit
Climate variables
yCO2
500
1600
ppm
yT
10
20
∘ C
yRH
0
80
%
Control inputs
uCO2
0
1.2
mg/m 2 /s
Table 4: Climate variable ranges and control input limits.
Hyperparameter
Value
Total training timesteps
4×106
Number of parallel environments ( nenvs )
8
Rollout length ( nsteps )
1920
Batch size
1920
Optimization epochs per update
10
Discount factor ( γ )
0.98
Table 5: PPO hyperparameters used in the experiments.
Hyperparameter
Symbol
Value
Initial penalty coefficient
λ0
1.0
Minimum penalty coefficient
λmin
0.05
Maximum penalty coefficient
λmax
50.0
Constraint threshold
d
0.7
Penalty update interval (episodes)
K
4
Penalty learning rate
αλ
0.01
Table 6: RCPO hyperparameters for constraint regulation.
Method
Cumulative profit
Δ Profit (%)
Cumulative violation
Δ Violation (%)
PPO (baseline)
3.782 ( ±0.146 )
– ( ±3.86 )
0.858 ( ±0.027 )
– ( ±3.15 )
RCPO w/o KAN & Time
3.543 ( ±0.061 )
−6.32 ( ±1.72 )
0.695 ( ±0.026 )
−19.00 ( ±3.74 )
RCPO w/o Time
3.635 ( ±0.064 )
−3.89 ( ±1.76 )
0.712 ( ±0.039 )
−17.02 ( ±5.48 )
RCPO w/o KAN
3.755 ( ±0.165 )
−0.71 ( ±4.39 )
0.710 ( ±0.025 )
−17.25 ( ±3.52 )
KAN-RCPO-PPO
3.892 ( ±0.147 )
+2.91 ( ±3.78 )
0.698 ( ±0.029 )
−18.65 ( ±4.15 )
Table 7: Performance comparison under different controller configurations.
Figure 3: Ablation study on policy representation enhancements. Training performance comparison between baseline PPO, PPO with cyclic time features, and PPO with KAN-based policy networks. All methods are trained without RCPO constraint regulation.
Figure 4: Normalized cumulative climate constraint violations for different climate variables achieved by the proposed RCPO-PPO controller relative to the baseline PPO controller. The baseline PPO result is normalized to 100% .
Figure 5: Normalized cumulative operation costs associated with heating and CO2 input achieved by the proposed RCPO-PPO controller relative to the baseline PPO controller. The baseline PPO result is normalized to 100% .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Parameter
Value
Parameter
Value
p1
0.544
p11
7.50×10−6
p21
3.60×10−3
p2
2.65×10−7
p12
8.31
p22
9348
p3
53
p13
273.15
p23
8314
p4
3.55×10−9
p14
101325
p24
273.15
p5
5.11×10−6
p15
0.044
p25
17.4
p6
2.30×10−4
p16
3.00×104
p26
239
Appendix
Table 8: Nominal parameter values of the lettuce greenhouse model.