Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to 220× and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.
Figures & tables
Figure 1: Overview of Jaxolotl ’s design and implementation. We provide a modular implementation of representative multi-task LTL-RL algorithms in terms of a generic task representation, policy architecture, and shared common modules. Precompilation of dynamic structures and environment resets enables end-to-end JIT compilation of the entire training and evaluation loop.
Figure 2: Visualisations of the main environments included in Jaxolotl .
Figure 3: End-to-end training and evaluation times of Jaxolotl compared to original reference implementations on ZoneEnv. We achieve speedups of up to 220× for training and 79× for evaluation.
Figure 4: Validation of Jaxolotl . (Left) Final ZoneEnv success rates are within ±1.1% of the reference implementations for all methods. (Right) Overall training dynamics of DeepLTL and GenZ-LTL closely follow the references.
Figure 5: Aggregate finite-horizon success rates over training. Lines show means across 10 independently trained policies, with shaded 95% Student- t confidence intervals.
LetterWorld
ZoneEnv
FrankaZoneEnv
Warehouse
Method
Fin. ↑
Inf. ↑
Fin. ↑
Inf. ↑
R-S ↑
Fin. ↑
Inf. ↑
Fin. ↑
Inf. ↑
R-S ↑
LTL2Action
0.42 ±0.03
—
0.43 ±0.03
—
—
0.04 ±0.01
—
0.34 ±0.09
—
—
GCRL-LTL
0.94 ±0.00
8.0 ±0.1
0.93 ±0.00
5.7 ±0.1
30.2 ±3.3
0.55 ±0.27
8.1 ±5.9
—
—
—
DeepLTL
0.87 ±0.01
5.3 ±0.2
0.92 ±0.02
3.0 ±0.3
540.1 ±48.0
0.51 ±0.17
4.8 ±3.3
0.80 ±0.09
2.6 ±0.3
618.1 ±118.7
GenZ-LTL †
0.99 ±0.00
8.2 ±0.0
1.00 ±0.00
6.5 ±0.2
492.0 ±50.6
1.00 ±0.00
35.0 ±0.6
—
—
—
↪ unreduced obs.
0.93 ±0.00
5.6 ±0.0
0.96 ±0.01
5.0 ±0.1
321.6 ±59.2
0.45 ±0.22
5.1 ±3.6
0.84 ±0.12
2.2 ±0.7
102.6 ±24.8
Table 1: Aggregated benchmark results over the Jaxolotl task suites, averaged over 512 episodes per formula, with 95% CIs over 10 training seeds. Fin.: finite-horizon success rate (0–1); Inf.: completed accepting cycles with fixed episode length. R-S: completed accepting cycles on the reach-stay task suite. Best results are bold , second best underlined , — denotes unsupported configurations. See Section H.1 for per-specification results.
Figure 6: (Left) Non-myopic reasoning results on different step variants of ConveyorWorldSimple- k . (Middle) Aggregate finite-horizon success rates on ZoneEnv-NM over training. (Right) Results of scaling the number of atomic propositions in FrankaZoneEnv.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Environment throughput of Jaxolotl and the reference implementations against the number of parallel environments (log scales). The references run at most 128 parallel environments; the dashed line marks their best measured throughput. Annotations give the peak throughput of Jaxolotl and its speedup over this best reference measurement.
Figure 8: (a) A visualisation of ConveyorWorld- k ; the green circle marks the agent start position, and the upper and lower one-way conveyor belts lead to rooms A and B respectively. (b) Non-myopic reasoning results on different step variants of ConveyorWorld- k .
ID
LTL Formula
LetterWorld
φ1
F(a∧(¬bUc))∧Fd
φ2
(Fd)∧(¬fU(d∧Fb))
φ3
(F((a∨c∨j)∧Fb))∧(F(c∧Fd))∧Fk
φ4
¬aU(b∧(¬cU(d∧(¬eUf))))
φ5
G¬(k∨l)∧F(a∧F(e∧F(i∧Fd)))
φ6
G¬(a∨b∨c∨d)∧G(h→¬fUg)∧Fg∧F(h∧Fi)
Appendix
Table 2: Finite-horizon LTL specifications used for evaluation.
ID
LTL Formula
LetterWorld
ψ1
GFa∧GFb∧GFc
ψ2
GF(a∧Fd)∧GF(b∧Fe)
ψ3
G¬l∧GFc∧GFg∧GFj
ψ4
(¬aUb)∧GFa∧GFf∧GFg∧GFh
ψ5
G(a→Fd)∧GFa∧GFh
ψ6
GF(c∧(¬dUe))∧GFb
Appendix
Table 3: Infinite-horizon LTL specifications used for evaluation.
ID
LTL Formula
ZoneEnv
χ1
FGred
χ2
FGred∧F(yellow∧Fgreen)
χ3
FGpurple∧G¬yellow
χ4
G((green∨yellow)→Fred)∧FG(green∨purple)
χ5
FGgreen
χ6
FGyellow
Appendix
Table 4: Reach-stay LTL specifications used for evaluation.
φ
LTL2Action
GCRL-LTL
DeepLTL
GenZ-LTL †
GenZ-LTL (unreduced obs.)
SemLTL
StructLTL
LetterWorld
φ1
0.62 ±0.14
0.99 ±0.00
1.00 ±0.00
0.99 ±0.00
1.00 ±0.00
1.00 ±0.00
1.00 ±0.00
φ2
0.74 ±0.15
0.99 ±0.00
0.96 ±0.01
1.00 ±0.00
0.98 ±0.00
0.95 ±0.01
0.97 ±0.01
φ3
0.45 ±0.13
1.00 ±0.00
1.00 ±0.00
1.00 ±0.00
1.00 ±0.00
1.00 ±0.00
1.00 ±0.00
φ4
0.52 ±0.20
0.98 ±0.01
0.95 ±0.01
0.99 ±0.00
0.97 ±0.01
0.94 ±0.01
0.95 ±0.01
φ5
0.02 ±0.01
0.92 ±0.01
0.73 ±0.03
1.00 ±0.00
0.85 ±0.02
0.57 ±0.07
0.79 ±0.04
φ6
0.00 ±0.00
0.72 ±0.01
0.53 ±0.07
0.97 ±0.01
0.77 ±0.02
0.32 ±0.04
0.62 ±0.03
Appendix
Table 5: Per-specification finite-horizon benchmark results (success rate), averaged over 512 evaluation episodes, with 95% Student- t CIs over training seeds. Best results are bold , second best underlined , — denotes unsupported configurations.
ψ
GCRL-LTL
DeepLTL
GenZ-LTL †
GenZ-LTL (unreduced obs.)
SemLTL
StructLTL
LetterWorld
ψ1
9.4 ±0.1
6.8 ±0.3
9.8 ±0.1
6.9 ±0.1
6.0 ±0.2
6.7 ±0.4
ψ2
6.4 ±0.0
4.3 ±0.2
6.7 ±0.1
4.5 ±0.0
4.1 ±0.1
4.3 ±0.2
ψ3
8.8 ±0.5
5.3 ±0.4
9.8 ±0.1
6.2 ±0.1
4.5 ±0.3
5.7 ±0.3
ψ4
6.3 ±0.0
4.0 ±0.3
6.7 ±0.1
4.5 ±0.0
3.3 ±0.2
4.1 ±0.2
ψ5
9.4 ±0.3
6.3 ±0.9
6.4 ±0.3
5.2 ±0.3
6.0 ±0.3
7.1 ±0.8
ψ6
9.3 ±0.1
6.4 ±0.5
9.9 ±0.1
6.8 ±0.2
4.2 ±0.6
6.8 ±0.4
Appendix
Table 6: Per-specification infinite-horizon benchmark results (completed accepting cycles), averaged over 512 evaluation episodes, with 95% Student- t CIs over training seeds. Best results are bold , second best underlined , — denotes unsupported configurations.
χ
GCRL-LTL
DeepLTL
GenZ-LTL †
GenZ-LTL (unreduced obs.)
SemLTL
StructLTL
ZoneEnv
χ1
24.2 ±5.0
645.1 ±85.6
548.0 ±57.2
340.6 ±119.3
615.7 ±38.5
608.4 ±61.8
χ2
51.2 ±16.2
497.2 ±61.5
471.7 ±61.7
307.4 ±99.8
445.5 ±22.1
501.2 ±56.8
χ3
21.6 ±3.3
512.8 ±79.5
533.1 ±51.9
317.6 ±103.0
591.4 ±67.9
575.4 ±63.8
χ4
28.4 ±6.7
520.9 ±85.7
495.6 ±37.5
303.1 ±104.9
285.3 ±101.4
571.0 ±64.9
χ5
25.9 ±4.3
635.4 ±51.3
541.4 ±51.2
333.8 ±129.0
679.1 ±64.9
593.8 ±100.8
χ6
18.9 ±2.6
513.4 ±113.3
299.4 ±94.5
341.8 ±104.1
639.0 ±66.7
626.8 ±97.9
Appendix
Table 7: Per-specification reach-stay benchmark results (completed accepting cycles), averaged over 512 evaluation episodes, with 95% Student- t CIs over training seeds. Best results are bold , second best underlined , — denotes unsupported configurations.
Category
Hyperparameter
LTL2Action
SemLTL
DeepLTL
StructLTL
GCRL-LTL
GenZ-LTL
PPO
Total environment steps
2×107
Environments
16
Steps/update
128
Minibatches
8
Update epochs
8
Discount ( γ )
0.94
Appendix
Table 8: Training and model hyperparameters for LetterWorld. Values spanning multiple columns are shared by those methods; N/A marks settings that do not apply to a method.
Category
Hyperparameter
LTL2Action
SemLTL
DeepLTL
StructLTL
GCRL-LTL
GenZ-LTL
PPO
Total environment steps
2×107
Environments
16
Steps/update
4,096
Minibatches
32
Update epochs
10
Discount ( γ )
0.998
Appendix
Table 9: Training and model hyperparameters for ZoneEnv and ZoneEnv-NM. Values spanning multiple columns are shared by those methods; N/A marks settings that do not apply to a method.
Category
Hyperparameter
LTL2Action
SemLTL
DeepLTL
StructLTL
GCRL-LTL
GenZ-LTL
PPO
Total environment steps
2×108
Environments
2,048
Steps/update
64
Minibatches
4
Update epochs
5
Discount ( γ )
0.995
Appendix
Table 10: Training and model hyperparameters for FrankaZoneEnv. Values spanning multiple columns are shared by those methods; N/A marks settings that do not apply to a method. The table covers FrankaZoneEnv- k for k∈{8,9,10,11,12} . Hyperparameters that vary with k are given as functions of k .
Category
Hyperparameter
LTL2Action
SemLTL
DeepLTL
StructLTL
GenZ-LTL
PPO
Total environment steps
2×108
Environments
1,024
Steps/update
128
Minibatches
4
Update epochs
5
Discount ( γ )
0.998
Appendix
Table 11: Training and model hyperparameters for Warehouse. Values spanning multiple columns are shared by those methods; N/A marks settings that do not apply to a method.
Category
Hyperparameter
LTL2Action
SemLTL
DeepLTL
StructLTL
GCRL-LTL
GenZ-LTL
PPO
Total environment steps
2×106
Environments
256
Steps/update
64
Minibatches
4
Update epochs
8
Discount ( γ )
0.9
Appendix
Table 12: Training and model hyperparameters for ConveyorWorld and ConveyorWorldSimple. Values spanning multiple columns are shared by those methods; N/A marks settings that do not apply to a method. The table covers ConveyorWorld- k for k∈{1,2,3,4,5,6,7,8,9,10} and ConveyorWorldSimple- k for k∈{1,2,4,8,16,32} . Hyperparameters that vary with k are given as functions of k . Note that environment resets are not precomputed for these environments (since they are trivially cheap), and the training “curriculum” only consists of a single stage (i.e. no curriculum) for all methods.