Interval timing is extensively studied as an important aspect of human behaviour. As artificial agents are increasingly designed to function alongside humans, their interval timing abilities also needs to be studied. However, research in this area remains limited and scattered. This paper presents Chronocooked, a reinforcement learning (RL) benchmark environment that enables a systematic study of interval timing abilities in RL agents. Inspired by Overcooked, the suite comprises cooking scenarios involving interval timing tasks drawn from the psychology literature. The tasks and reward functions are designed such that temporal information is unobserved but critical for optimal performance. The environment is intentionally kept simple to enable controlled experiments and support biologically plausible models. Each task is accompanied by evaluation metrics to study different aspects of interval timing in RL agents, namely, task performance, human-like timing and scalability. We report baselines using a non-recurrent architecture (CNN), a recurrent architecture (LSTM), and a biologically inspired recurrent architecture (CTRNN). The baseline model analysis shows that, although RL agents can successfully perform time-dependent tasks, they do not necessarily process and perceive time in the same way as humans. Understanding these differences is important for anticipating their impact on human-robot interactions (HRI).
Figures & tables
Figure 1 . Chronocooked grid layout for the two core tasks. All other tasks discussed in this study are modifications of these two tasks to evaluate different temporal processing abilities of an RL agent. The Chronocooked grid layout consisting of an onion dispenser(in pink), an oven(in yellow), a delivery counter (in orange or green) for the FI timing task and the temporal bisection task
Figure 2 . FI timing (task performance): First oven check (FOC) and number of timing policies (top and bottom respectively) across model types (shown in different colours) and TD (x-axis). The standard deviation is shown as error bars on corresponding barplots. Missing barplots indicate that models did not converge in the soup delivery task in any of the runs. In the top plot, the barplots with a yellow border indicate cases where mean FOC was not significantly different from the corresponding TD. The mean FOC of the high memory recurrent models especially LSTM models closely tracks the target duration. These models also exhibit variable action sequences in the timing phase. Together these results indicate a tendency of implicit time-keeping in high memory recurrent models
Figure 3 . FI timing (human-like timing): The plots show results for the three regularities of scalar property. The leftmost plot also shows evidence for FOC validity because the mean FOC increases linearly with TDs. However, none of the models follow all scalar property regularities. The corresponding regression statistics for each plot are shown in blue at the bottom. LSTM125, LSTM256 and CTRNN256 showed strong positive correlation between the SD and mean FOC. The CV (SD/mean) showed no significant relationship with TD for all models except LSTM256, consistent with the constant CV regularity expected under scalar timing. Taken together, while some models showed partial conformity to individual regularities of the scalar property, no model satisfied all three regularities simultaneously.
Figure 4 . Multi-timer categorical variant (task performance): The plots show the average reward of the different models for two TD configurations (row titles) in the last two CL stages, namely, 4-timer and 5-timer setting (column titles) for the multi-timer task with buffer. The error bars depict the SD across the different model runs and seeds. The red dotted line indicates the CL performance threshold. All CTRNN models perform better than the LSTM models in both TD configurations. LSTM8 shows the worst performance in the 5-timer setting.
Figure 5 . Temporal bisection task (task performance): Psychometric curves of the three model architectures (rows) for the different anchors (columns). Each plot shows P(long) plotted against the different probe durations. The arithmetic mean (AM) and geometric mean (GM) of each anchor are shown at the bottom of each column. The grey dots represent p(long) of each run and the solid grey line shows the corresponding logistic fit. Similarly, the black dots represent the average P(long) across all runs and the red line shows the logistic fit on the average P(long). The vertical blue dotted lines corresponds to durations with P(long) = 25% and P(long) = 75% respectively. The region between the two blue dotted lines is called the Just Noticeable Difference (JND). Smaller JND corresponds to higher discriminability and vice versa. The bisection points (BPs) (duration at which P(long) = 50%) are shown in the respective plots. The RL agents with sufficient memory and a recurrent architecture can both converge on the soup delivery task and perform the bisection task including OOD. They also exhibit an s-shaped response curve which is qualitatively similar to human and animal data.
Figure 6 . Temporal bisection task (Human-like timing): (a) The plots show three BP regularities (y-axis) as a function of spread (x-axis). Both architectures show that BP tracks GM and stays around 1.2xGM even at larger spreads, unlike humans, where BP track the AM. (b) The plots show Weber ratio against spread (first column) and against AM for spreads of 2 and 3 (second and third column, respectively). Both architectures show decreasing discriminablity with spread, consistent with human data. For a specific spread, both architectures show a stronger positive trend in Weber ratio than human data (r=0.02) and CTRNN violates Weber’s law at spread 3. (a-b) Each plot shows the corresponding statistics of the regression fits,namely, the Pearson correlation coefficient (r), p-value (p) and number of samples (n). The results for corresponding human data taken from Kopec and Brody [2010] are shown at the bottom of each columns. The BP plots show that for both architectures, BP tracks GM and stays around 1.2xGM even at larger spreads, unlike humans, where BP track the AM. The weber ratio plots show that, both architectures exhibit decreasing discriminablity with spread, consistent with human data. For a specific spread, both architectures show a stronger positive trend in Weber ratio than human data (r=0.02) and CTRNN violates Weber’s law at spread 3
Effect
LSTM
CTRNN
Comments (contrast from human data)
BP: sub-geometric → sub-arithmetic
×
–
Opposite direction (both models).
BP: BP/AM vs log2(L/S) (BP track AM)
×
×
Same trend as humans, but with a steeper decrease (both models)
BP: sub-arithmetic for small spreads
×
✓
However, CTRNN declines with spread rather than remaining at a constant fraction of AM.
BP: BP/GM vs log2(L/S)
×
×
Same trend as humans, but with a less steep increase (both models)
BP: supra-geometric for high spreads
×
×
BP remains below 1.2 times GM (both models)
Weber ratio (discriminability)
✓
✓
Same trend as humans, but with a steeper increase (both models)
Table 2 . Temporal bisection task (Human-like timing): Summary of BP and Weber ratio regularities relative to human data.
Figure 7 . Multi-Anchor categorization Task in Chronocooked. The example shows 4 delivery counters (in different colours) each for a different anchor duration. The grid layout for Multi-Anchor categorization Task in Chronocooked. It is similar to the temporal bisection task layout except that it contains more delivery counters. The example shows 4 delivery counters (in different colours) each for a different anchor duration.
Figure 8 . Multi-anchor (Temporal categorization ability): Visualization of 3-anchor (left) and 4-anchor(right) setting. Each subplot shows a stacked bar plot indicating the percentage of deliveries in each delivery counter (indicated by a different colour), aggregated across four runs for a given RL agent (rows) and anchor configuration (columns). Delivery counter 1 corresponds to the smallest anchor duration, delivery counter 2 to the second smallest, and so on. A smooth transition from delivery counter 1 to delivery counter 2 and then 3 and 4 indicates a good performing model. Grey bars denote failed cases, in which the agent did not deliver soup to any counter. The figure shows a clear transition from delivery counter 1 to 2, then 3, and then 4, as probe durations move from one extreme to the other. Importantly, this includes OOD probe durations. CTRNN models exhibit better performance as compared to LSTM models in the 4-anchor setting (last CL stage).
Figure 9 . Multi-anchor (Task performance): The left panel shows the model-level task performance (as described in Section 8.1) of each recurrent model. Specifically, the plots show the percentage of incorrect deliveries aggregated across runs and main anchor configurations for the different n-anchor settings (indicated by different colours). The red stars on the x-axis denote probe duration categories where % of incorrect deliveries in the 4-anchor setting was significantly worse than the corresponding 3-anchor setting. The Wilcoxon signed-rank test was used for the comparison. The right panel shows the architecture-level comparison for the 4-anchor setting (last CL stage) task performance. Data across the different main anchor configurations and the two memory sizes within each architecture is combined. The red stars on x-axis in this case denote the probe duration category where LSTM task performance was significantly worse than CTRNN. A Mann-Whitney U test was used for the comparison. For both plots, the error bars show the SD and the p-values are corrected using the Holm–Bonferroni method. Refer Appendix B for details on all significance tests. CTRNN models showed a better task performance than LSTM models both for the model-level and architecture-level analysis.
Figure 10 . Asynchronous Multi-Oven Bisection Task in Chronocooked. The example shows 3 oven each with a different start time and different TD. The example shows the grid layout for multi-oven bisection task. The layout is similar to the temporal bisection task, except that there are 3 ovens in the last row instead of just 1 oven.
Figure 11 . Asynchronous Multi-Oven Bisection (task performance): The plots show the success rate (successful bisection tasks) across all ovens, probe durations (including OOD durations) and runs, for each model (shown in different colours) and anchor configuration (columns), in 2-oven and 3-oven setting (x-axis). While the asynchronous multi-oven task was in general difficult, performance improved as distance between S and L anchor increased. In the asynchronous variant, performance improved as distance between S and L increased. In the 3-oven setting, models achieved at most a 50% success rate for the 4–18 anchor configuration (maximum distance between S and L), while performance was lower for all other anchor configurations.
Figure 12 . Asynchronous multi-oven bisection (Scalability): The plot shows the training time steps required for each CL stage (n-oven setting) shown in different colours for each model (rows) and anchor configurations (columns). The numbers in the bars represent the corresponding CL-efficiency ratios. All models needed significantly more than 1x the time steps in 1-oven setting to converge in the 2-oven setting (CL-efficiency ratio >1). All models reach the 3-oven setting only for anchor configuration 4-18 (largest spread). Only one run from LSTM125 reaches 3-oven setting for anchor configuration 4-8 (smallest spread). All models need significantly more than 1x time steps of previous CL stage to complete the current CL stage.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 13 . Multi-timer exact variant (task performance): The plots show the average reward of the different models (x-axis) in the different CL stages (columns) for the 2 TD configurations (rows) of the multi-timing exact-timer variant. CTRNN models fail already in the 1-timer stage. The LSTM models perform better in anchor configuration with smaller distance between consecutive anchors. LSTM8 reached up to 5-timer stage surpassing the larger memory models. This could be attributed to a fix timing policy followed by this model.
Figure 14 . Multi-timer, exact-timer variant (# timing policies): The plots show the number of unique timing policies in the different seeds, averaged across runs, for the different RL agents (shown in different colours) for the anchor configuration 2-4-6-8-10 in the exact-timer variant. The plot titles indicate the different CL stages. As in the FI timing task, LSTM8 showed a lower number of timing action sequences compared to the two higher memory LSTM models (see Figure \ref{fig:multitimer_num_timing_actions_exact}), suggesting greater reliance on embodied timing (a fixed action sequence) rather than implicit time-keeping. This could be one of the possible reasons for LSTM8 model to successfully reach 5-timer setting in the more difficult exact-timer variant of the multi-timer task.
Figure 15 . Multi-timer categorical variant (CL-efficiency): The training time steps required by the different runs of the the models (rows) for the TD configurations (columns), split by CL stages (shown in different colours) for categorical variant. The numbers in the bars represent the corresponding CL-efficiency ratios. Except for CNN, at least one run from all models reach the last CL stage (5-timer) in the categorical timer variant
Figure 16 . Multi-timer exact-timer variant (CL-efficiency): The training time steps required by the different runs of the the models (rows) for the TD configurations (columns), split by CL stages (shown in different colours) for exact-timer variant. The numbers in the bars represent the corresponding CL-efficiency ratios. All runs in the CTRNN models fail already in the 1-timer stage. The LSTM models perform better in anchor configuration with smaller distance between consecutive anchors. 2 runs from LSTM8 reached up to 5-timer stage surpassing the larger memory models. This could be attributed to a fix timing policy followed by this model
Figure 17 . Multi-anchor categorization task performance: The plots show the percentage of incorrect deliveries across runs and seeds for a given probe duration category, with error bars indicating +- standard deviation. The final column shows the mean and standard deviation across runs and anchor configurations. Lower values indicate better performance; Models should be compared along each probe-duration interval (<=A1, A1–A2, etc.) to isolate the effect of probe durations. The red stars in the average column indicates the cases when anchor 4 performance was significantly worse than corresponding anchor 3 performance. At an architecture level, for 4 anchor setting, CTRNN yielded significantly fewer incorrect deliveries than LSTM in 3 of the 5 probe categories: Probe category 0 (F = 14.93, p < .001), Probe category 1 (F = 4.48, p = .041), and Probe category 4 (F = 9.27, p = .004) At an architecture level, for 4 anchor setting, CTRNN yielded significantly fewer incorrect deliveries than LSTM in 3 of the 5 probe categories
Figure 18 . Multi-anchor (CL efficiency): Training time steps required by each run of the models (rows) for the different CL stages (shown in different colours) across all main anchor configurations (columns). The numbers in the bars represent the corresponding CL-efficiency ratios. The last column shows an average taken across the main anchor configurations for each model. Most runs from all models reached the last CL stage (4-anchor) across all anchor configurations.
Model
Probe category
N
4-anchor Median [IQR]
3-anchor Median [IQR]
W
p
pHolm
Reject H0
LSTM125
<=A1
15
0.19 [0.23]
0.00 [0.16]
76.00
0.016
0.033
Yes
LSTM125
A1-A2
15
0.11 [0.42]
0.00 [0.01]
75.00
0.002
0.007
Yes
LSTM125
A2-A3
15
0.15 [0.58]
0.00 [0.00]
102.00
< .001
0.004
Yes
LSTM125
A3-A4
15
0.09 [0.33]
0.00 [0.10]
82.00
0.032
0.033
Yes
LSTM256
<=A1
12
0.27 [0.69]
0.05 [0.10]
43.00
0.006
0.023
Yes
LSTM256
A1-A2
13
0.08 [0.19]
0.02 [0.04]
49.00
0.087
0.087
No
Appendix
Table 3 . Multi-anchor (Task performance): Wilcoxon signed-rank comparisons between the 3-anchor and 4-anchor conditions. p -values were corrected using the Holm method.
Model
Probe category
N 125 / 256
Memory=125 Median [IQR]
Memory=256 Median [IQR]
Statistic
p -value
pHolm
Reject H0
LSTM
<=A1
15 / 12
0.19 [0.23]
0.27 [0.69]
76.50
0.524
1.000
No
LSTM
A1-A2
15 / 13
0.11 [0.42]
0.08 [0.19]
106.50
0.689
1.000
No
LSTM
A2-A3
15 / 12
0.15 [0.58]
0.18 [0.40]
84.00
0.788
1.000
No
LSTM
A3-A4
15 / 12
0.09 [0.33]
0.53 [0.87]
79.00
0.606
1.000
No
LSTM
>=A4
15 / 11
0.59 [0.86]
1.00 [0.51]
67.00
0.418
1.000
No
CTRNN
<=A1
13 / 11
0.00 [0.00]
0.00 [0.18]
52.50
0.159
0.648
No
Appendix
Table 4 . Multi-anchor (task performance): Mann–Whitney U test comparing the percentage incorrect deliveries in 4-anchor setting between the two memory sizes (125 and 256) of a specific model architecture.
Figure 19 . Multi-oven bisection (task performance): The top plot depicts the asynchronous task performance and bottom plot shows the synchronous task performance. The plots show the success rate (successful bisection tasks) across all ovens, probe durations (including OOD durations) and runs, for each model (shown in different colours) and anchor configuration (columns), in 2-oven and 3-oven setting (x-axis). While the asynchronous multi-oven task was in general difficult, performance improved as distance between S and L anchor increased. All the RL agents achieved good performance in the synchronous variant of the task. In the asynchronous variant, performance improved as distance between S and L increased. In the 3-oven setting, models achieved at most a 50% success rate for the 4–18 anchor configuration (maximum distance between S and L), while performance was lower for all other anchor configurations.
Figure 20 . Asynchronous multi-oven bisection (psychometric curves): The plot shows the psychometric curve of each model (rows) for the different ovens in the 2-oven (left) and 3-oven setting (right), for the anchor configuration 4-18 which was the best performing anchor (also the one with the largest distance between the anchors). The primary oven (oven 1) shows the best performance among the different ovens, especially in the 3-oven setting. The primary oven (oven 1) shows the best performance among the different ovens, especially in the 3-oven setting.
Figure 21 . Asynchronous Multi-oven bisection (psychometric curve): The plot shows the psychometric curve for the different ovens for 2-oven and 3-oven setting for anchor 4-12 (top) and 4-8(bottom). All model performance degrades in the 3-oven setting. The primary oven (oven 1) shows the best performance among the different ovens. The performance in the 3-oven stage for the asynchronous timing task is better in the 4-12 anchor configuration rather than the 4-8 configuration. All model performance degrades in the 3-oven setting. The primary oven (oven 1) shows the best performance among the different ovens.
Figure 22 . Synchronous multi-oven bisection (scalability): The plots show the training time steps required for each CL stage (n-oven setting) shown in different colours for each model (rows) and anchor configuration (columns) for the synchronous variant. The numbers in the bars depict CL-efficiency ratios. At an architecture level, the CL-efficiency ratio of both architectures was significantly more than 1. In the synchronous variant most model runs reached the last CL stage (3-oven).
TD
Model
N
M
SD
t
p
4
LSTM8
36
3.39
0.87
-4.21
0.000
4
LSTM125
48
3.98
0.53
-0.27
0.785
4
LSTM256
48
3.96
0.68
-0.42
0.674
4
CTRNN8
24
0.00
0.00
-inf
0.000
4
CTRNN125
42
2.52
1.44
-6.66
0.000
4
CTRNN256
36
1.97
1.63
-7.46
0.000
Appendix
Table 5 . FI timing: Significance test comparing FOC with the corresponding TD for each model
Figure 23 . Psychometric curves of the models (rows) for the different anchors (columns). Y-axis shows percentage of long - P(long) and x-axis shows probe (test) durations along the the arithmetic mean (AM) and geometric mean (GM) of each anchor. The grey dots represent p(long) of each run (with 11 seeds each). The black dots is the average P(long) across all runs. The red line shows the sigmoid fit on the average P(long). The bisection points (BP) are shown in the respective plots. The smaller memory sizes in the recurrent models do not perform as good as the corresponding high memory models