GUIDE: Reinforcement Learning for Behavioral Action Support in Type 1 Diabetes
Authors: Saman Khamesian, Sri Harini Balaji, Di Yang Shi, Stephanie M. Carpenter, Daniel E. Rivera, W. Bradley Knox, Peter Stone, Hassan Ghasemzadeh
Organizations: College of Health Solutions, Arizona State University, Phoenix, USA · School of Computing and Augmented Intelligence, Arizona State University, Tempe, USA · School for Engineering of Matter Transport and Energy, Arizona State University, Tempe, USA · Institute for Foundation of Machine Learning, The University of Texas at Austin, Austin, USA · Sony AI, Austin, USA
Type 1 diabetes (T1D) management requires continuous adjustment of insulin and lifestyle behaviors to maintain blood glucose within a safe target range. Although automated insulin delivery (AID) systems have improved glycemic outcomes, many patients still fail to achieve recommended clinical targets. Current reinforcement learning (RL)-based methods focus primarily on insulin-only treatment and do not provide behavioral recommendations for glucose control. To address this gap, we propose GUIDE, an RL-based decision-support framework designed to complement AID technologies by providing structured behavioral recommendations defined by intervention type, magnitude, and timing, including bolus insulin administration and carbohydrate intake events. GUIDE integrates a patient-specific glucose predictor trained on real-world continuous glucose monitoring data and supports offline and online RL algorithms within a unified environment. The algorithms are evaluated across 25 individuals from the AZT1D dataset and 12 individuals from the OhioT1DM dataset. Among the evaluated algorithms, CQL-BC achieved the best performance on both datasets, with mean time-in-range values of 84.18 ± 19.89% on AZT1D and 77.64 ± 8.81% on OhioT1DM. It also maintained time-below-range values of 0.43 ± 1.27% and 2.81 ± 3.56%, respectively. Behavioral analysis yielded mean cosine similarities of 0.767 ± 0.128 on AZT1D and 0.774 ± 0.162 on OhioT1DM, indicating that the learned policy preserves key structural characteristics of patient action patterns. These findings demonstrate the potential of conservative offline RL with a structured behavioral action space to provide personalized and behaviorally plausible decision support for diabetes management.
Figures & tables
Fig. 1: Illustration of an artificial pancreas system with closed-loop blood glucose control in patients with T1D. Glucose measurements from a continuous glucose monitor are processed by a control algorithm to compute insulin dosing delivered through an insulin pump.
Fig. 2: Schematic overview illustrating the complementary roles of the AID system and the GUIDE framework, with both components utilizing glucose measurements to inform insulin delivery and behavioral action recommendations for patients with T1D.
Fig. 3: Overview of the GUIDE framework. Data are partitioned chronologically within each subject, with 80% used to train the glucose level predictor (GLIMMER) and 20% reserved for defining initial states. The trained predictor is deployed within the environment simulator alongside the human-inspired meal generator and glucose-responsive basal insulin generator. At each decision step, the RL agent selects a behavioral action based on the current state, and the environment returns the resulting reward and next state, forming a closed-loop learning process.
Fig. 4: Historical basal-rate distribution and basal generator function. Blue contours show the relative density of hourly mean CGM–basal-rate pairs, orange markers and error bars show the median and interquartile range, and the red line represents the generator function. Black and gray dashed lines mark the clinical glycemic thresholds (70 and 180 mg/dL) and clinical cut-off thresholds (100 and 250 mg/dL), respectively.
Meal
Time Window
CHO Distribution (g)
Breakfast
07:00–09:00
TN(65,152,20,100)
Lunch
12:00–14:00
TN(65,152,20,100)
Dinner
19:00–22:00
TN(65,152,20,100)
TABLE I: Meal occurrence windows and CHO (carbohydrate) intake parameters used by the human-inspired meal generator. For each meal, a start hour is randomly selected from the specified time range. Carbohydrate portions are sampled from a truncated normal distribution (mean = 65 g, SD = 15 g, range = 20–100 g).
Metric
Description
Target
TIR
Percentage of time glucose is within 70–180 mg/dL
>70%
TBR
Percentage of time glucose is below 70 mg/dL
<4%
TAR
Percentage of time glucose exceeds 180 mg/dL
<25%
CV
Glucose variability
<36%
TABLE II: Glycemic control metrics used for evaluation with clinically recommended targets [ 4 , 40 ] .
Dataset
RMSE (mg/dL)
MAE (mg/dL)
Dysglycemia F1 (%)
AZT1D
22.48±3.57
15.58±2.87
82.07
OhioT1DM
23.97±3.77
15.83±2.09
91.60
TABLE III: Prediction performance of GLIMMER on the AZT1D and OhioT1DM datasets using a 60-minute prediction horizon. The complete evaluation results are reported in [ 31 ] .
Fig. 5: GLIMMER counterfactual glucose trajectories under increasing bolus-insulin and carbohydrate inputs. For each seed, predictions were averaged across the 10 test states within each patient and subsequently across the 25 patients. Solid lines represent the mean of the five seeds, and shaded regions indicate ±1 SD across seeds.
Intervention
Level
Mean ΔG (mg/dL)
ΔG60 (mg/dL)
Insulin
2 U
−0.24±4.52
−0.35±4.07
Insulin
5 U
−1.49±9.96
−1.65±8.97
Insulin
10 U
−5.07±15.00
−5.23±13.74
Carbohydrate
10 g
1.91±2.08
1.55±2.14
Carbohydrate
30 g
4.77±6.36
3.78±6.38
Carbohydrate
50 g
6.09±9.73
4.79±9.71
TABLE IV: Counterfactual response magnitude relative to the matched zero-action prediction. Values are cohort mean ± SD across patient-level averages after averaging the five seeds.
Expected ordering
Directional
Complete ranking
0U>2U>5U>10U
82.9%±3.7%
71.3%±4.8%
0g<10g<30g<50g
85.4%±4.6%
77.7%±3.3%
TABLE V: Directional and complete-ranking consistency of GLIMMER counterfactual responses (mean ± SD across five seeds). Directional consistency is calculated from 750 nonzero-versus-zero comparisons per seed, and complete-ranking consistency from 250 patient–state evaluations per seed.
H (h)
Hourly Error (mg/dL)
Cumulative Error (mg/dL)
MAE
RMSE
MAE
RMSE
1
14.16±3.25
21.20±4.76
14.16±3.25
21.20±4.76
6
20.56±5.47
29.22±6.85
18.48±4.50
27.37±5.84
12
21.45±6.12
30.76±7.41
19.41±5.05
28.49±6.58
18
22.50±6.71
31.50±8.77
20.77±6.10
29.92±7.44
24
22.79±8.36
32.90±10.46
20.91±8.56
31.12±9.68
TABLE VI: GLIMMER prediction error during the recursive 24-hour rollout at selected horizons H , where H denotes the rollout hour. Hourly metrics are calculated from the 12 predictions within hour H , whereas cumulative metrics include all predictions from hour 1 through H . Values are reported as mean ± standard deviation across 25 patients.
Fig. 6: Representative full-day simulation under full adherence using the personalized glucose prediction model. The solid green curve shows the predicted glucose trajectory across 24 decision steps. The dashed lines at 70 and 180 mg/dL mark the hypoglycemia and hyperglycemia thresholds, respectively, and bound the target range. Vertical markers denote decision and meal events: green indicates no intervention (Nothing), blue indicates carbohydrate intake recommended by the RL agent (Eat), red indicates a recommended bolus insulin dose (Inject), and magenta indicates a structured meal generated by the human-inspired meal controller. Numerical labels report carbohydrate amounts in grams for Eat and Meal events and insulin doses in units for Inject events.
Algorithm
TIR (%) ↑
TAR (%) ↓
TBR (%) ↓
CV (%) ↓
Policy Type
Interaction Mode
AZT1D Dataset
TD3-BC
81.64±20.59
17.83±20.69
0.53±1.27
15.49±3.74
off-policy
offline
CQL-BC
84.18±19.89
15.39±19.99
0.43±1.27
15.13±3.88
off-policy
offline
SAC-Offline
71.32±25.17
27.59±25.95
1.09±1.98
16.86±4.19
off-policy
offline
PPO
78.25±20.43
21.06±20.70
0.68±1.43
15.32±3.79
on-policy
online
SAC-Online
70.55±25.09
28.35±25.72
1.10±2.02
16.85±4.11
off-policy
online
TABLE VII: Performance comparison of RL algorithms within the proposed GUIDE framework under identical environment and reward settings across the AZT1D and OhioT1DM datasets. Results are reported for offline algorithms (TD3-BC, CQL-BC, SAC-Offline) and online algorithms (PPO, SAC-Online). Metrics TIR, TAR, TBR, and CV are reported as mean ± standard deviation across 25 subjects for AZT1D and 12 subjects for OhioT1DM. The Random policy illustrates outcomes under unstructured action selection within the simulation environment, while History summarizes each subject’s historical glycemic outcomes for contextual reference.
Algorithm
TD3-BC
CQL-BC
PPO
SAC-Offline
SAC-Online
Random
AZT1D Dataset
TD3-BC
–
4.20×10−2
6.48×10−3
8.94×10−7
2.15×10−6
1.23×10−4
CQL-BC
4.20×10−2
–
1.23×10−4
1.55×10−6
8.94×10−7
1.23×10−4
PPO
6.48×10−3
1.23×10−4
–
1.87×10−4
2.82×10−5
1.50×10−3
SAC-Offline
8.94×10−7
1.55×10−6
1.87×10−4
–
7.05×10−1
6.91×10−1
SAC-Online
2.15×10−6
8.94×10−7
2.82×10−5
7.05×10−1
–
7.05×10−1
TABLE VIII: Pairwise comparison of algorithms based on TIR 1 across the AZT1D and OhioT1DM datasets. Values represent Holm–Bonferroni adjusted p-values obtained from Wilcoxon signed-rank tests using per-subject TIR across 25 subjects for AZT1D and 12 subjects for OhioT1DM. The correction controls the family-wise error rate across the 15 pairwise comparisons within each dataset. Bold values denote statistically significant differences ( p<0.05 ) after adjustment.
Statistic
Cosine Similarity ↑
MRD ↓
dL1↓
AZT1D Dataset
Mean
0.767
0.794
0.667
STD
0.128
0.261
0.208
Median
0.779
0.799
0.663
OhioT1DM Dataset
Mean
0.774
0.840
0.712
TABLE IX: Aggregate behavioral similarity metrics for the CQL-BC algorithm across the AZT1D and OhioT1DM datasets. The AZT1D results include 24 subjects, with subject #23 excluded because of missing values, while the OhioT1DM results include all 12 subjects.
Configuration
TIR (%) ↑
TBR (%) ↓
TAR (%) ↓
Complete
84.40±19.70
0.50±4.97
15.10±19.85
Fixed timing
74.09±21.47
3.85±6.20
22.07±21.50
Carbohydrate only
71.00±21.26
1.49±5.12
27.51±21.50
Insulin only
72.30±20.01
5.69±5.56
22.01±20.17
Random
62.07±25.02
5.79±7.42
32.13±26.24
TABLE X: Glycemic performance under different action-space configurations. Results are reported as mean ± standard deviation across 25 patients, with 10 evaluation tests per patient.
Configuration
TIR (%) ↑
TBR (%) ↓
TAR (%) ↓
Complete
84.40±19.70
0.50±4.97
15.10±19.85
No glycemic
64.77±27.79
5.65±5.73
29.58±28.48
No meal
74.81±23.88
3.60±5.86
21.59±23.05
No insulin
73.84±22.54
1.90±6.29
24.26±22.15
TABLE XI: Glycemic performance under different reward function configurations. Results are reported as mean ± standard deviation across 25 patients, with 10 evaluation tests per patient.
Adherence target
TIR (%) ↑
TBR (%) ↓
TAR (%) ↓
100%
84.40±19.70
0.50±4.97
15.10±19.85
75%
78.19±22.68
2.63±5.35
19.18±21.94
50%
68.45±24.37
4.97±6.16
26.58±25.08
TABLE XII: Glycemic performance under different adherence targets. Results are reported as mean ± standard deviation across 25 patients, with 10 evaluation tests per patient.
Offline reinforcement learning (ORL) offers the potential to improve the quality of clinical decision-making using historical electronic health record (EHR) data. Current training and evaluative practices in this field rely heavily on EHR datasets that have been temporally discretised into fixed, regular time intervals. Discretisation creates fictional representations of complex clinical scenarios and compromises the generalisability of retrospective model evaluations. In this paper, we introduce Insulin4RL, a healthcare ORL dataset featuring naturally irregular inputs and actions from real clinical trajectories. Derived from MIMIC-IV, Insulin4RL comprises over 375,000 labelled decisions across 12,209 patients requiring insulin infusion titration in the Intensive Care Unit. The dataset can thus be used for research into ORL model performance under realistic clinical sampling assumptions. We provide a description of the dataset's structure and characteristics, baseline performance metrics using model-free offline reinforcement learning, and a standardised evaluation protocol using fitted Q-evaluation. We conclude with suggested areas for future research that could be addressed using this resource.
Thomas Frost, Steve Harris
Institute of Health Informatics University College London London, United Kingdom
Reinforcement learning (RL) in healthcare has had mixed results, with reward sparsity, unreliable off-policy evaluation, and deployment-simulation gap as recurring failure modes. We argue that chronic disease management is structurally a more tractable RL setting than the acute-care problems the field has primarily studied, but only if the problem is formalized to exploit chronic care's properties. We propose such a formalization. The agent's objective is to compress time-to-control (TTC) under a tiered reward calibrated to the CMS ACCESS Model. Two quantities from our companion preference-learning paper [Singh et al. 2026] enter as load-bearing structural elements: the execution intensity εbounds action availability under a constrained Markov Decision Process, and the clinician capability κweights offline-data transitions during RL training. Together they couple preference learning and RL into a two-loop architecture. We present simulation results on synthetic state machines for hypertension and type 2 diabetes. Capability-weighted offline RL outperforms uniform-weighted offline RL and the behavior policy by 15 percentage points on T2D TTC; the uniform-weighted formulation (the standard in existing healthcare RL) underperforms even the heterogeneous behavior policy. \Epsilon-aware policies generalize across deployment regimes while ε-naive policies do not.
Accurate long-horizon glucose forecasting is critical for automated insulin delivery systems, which help people with type 1 diabetes (T1D) manage their glucose and avoid dangerous hypoglycemia. However, standard recursive long short-term memory (LSTM) networks suffer from systematic negative bias at longer horizons due to error compounding, while purely mechanistic ordinary differential equation (ODE) models fail to generalize across individuals when parameterized at the population level. We propose PhysioSeq2Seq, a hybrid architecture that combines patient-specific physiological modeling with a sequence-to-sequence (Seq2Seq) LSTM. For each glucose segment, twin matching searches a population of 300 parameterized digital twins to identify the best-fitting physiological match from a 3-hour continuous glucose monitoring (CGM) history. The 10 internal ODE state variables of the matched twin are injected as exogenous covariates into both the encoder and decoder of the Seq2Seq LSTM. This simultaneous 48-step prediction strategy eliminates recursive error compounding, while the ODE features provide a physics-grounded constraint that bounds long-horizon drift within physiologically plausible ranges. PhysioSeq2Seq was trained on CGM and insulin data from 348 participants in the Type 1 Diabetes Exercise Initiative (T1DEXI) dataset and evaluated on 74 held-out participants. At the 240-minute horizon, PhysioSeq2Seq achieves a mean absolute error of 39.28 mg/dL and a mean error of -10.62 mg/dL, reducing bias by 13.89 mg/dL over the recursive LSTM and reducing mean absolute error by 28.62 mg/dL over the ODE-based digital twin. These results show that eliminating architectural feedback and injecting patient-matched physiological states is an effective and clinically meaningful strategy for long-horizon glucose forecasting in T1D.
Phat Tran, Neville Mehta, Clara Mosquera-Lopez +3
Oregon State University · Oregon Health & Science University