Meta-reinforcement learning with minimum attention
Organizations: Department of EECS University of Michigan Ann Arbor, MI 48109 · Department of Mathematics Morgan State University Baltimore, MD 21251
Abstract
Minimum attention applies the least action principle in changes of control concerning state and time, first proposed by Brockett. The involved regularization is highly relevant in emulating biological control, such as motor learning. We apply minimum attention in reinforcement learning (RL) as part of the rewards and investigate its connection to meta-learning and stabilization. Specifically, model-based meta-learning with minimum attention is explored in high-dimensional nonlinear dynamics. Ensemble-based model learning and gradient-based meta-policy learning are alternately performed. Empirically, minimum attention improves fast adaptation in few shots and reduces variance from perturbations of the model and environment, compared to model-free and model-based RL baseline, and yields consistent gain when integrated into modern world models (DreamerV3, MAMBA). Furthermore, the minimum attention demonstrates an improvement in energy efficiency.
Figures & tables
| MB-MPO | Ours ( ) | Ours ( ) | Ours ( ) | |
|---|---|---|---|---|
| Meta-training | ||||
| Average total rewards | 6692 318 | 9721 128 | 7545 123 | 8385 179 |
| Average feedback normalized | 783 135 | 567 53 | 590 88 | 637 56 |
| Average feedforward normalized | 5.16 0.11 | 5.01 0.07 | 5.02 0.10 | 5.05 0.07 |
| Average energy normalized | 4.54 0.13 | 3.85 0.05 | 3.94 0.03 | 3.91 0.02 |
| Meta-testing |
| Half-Cheetah | Hopper | Walker2D | Humanoid | |
| MB-MPO training | 6692 318 | 2475 3.81 | 2399 694 | 575 177 |
| MB-MPO meta-testing | 6356 132 | 467 100 | 523 174 | 315 67 |
| MB-MPO + Min Attn training | 9721 128 | 2825 1.25 | 3038 148 | 498 38 |
| MB-MPO + Min Attn meta-testing | 6822 109 | 485 61 | 1123 152 | 580 34 |
| HalfCheetah main results | |||
|---|---|---|---|
| Setting | Metric | MB-MPO | Ours (MA) |
| Meta-train | Reward | 6692 318 | 9721 128 |
| Feedback | 783 135 | 567 53 | |
| Energy | 4.54 0.13 | 3.85 0.05 | |
| Meta-test | Reward | 6356 132 | 6822 109 |
| Feedback | 673 75.6 | 625 34.7 | |
| Metric | Ours (Full MA) | Jacobian-only | Temporal-only |
|---|---|---|---|
| Reward | 9721 128 | 6484 148 | 6475 135 |
| Feedback | 567 53 | 128 24.5 | 736.2 84.9 |
| Feedforward | 5.01 0.07 | 9.54 0.35 | 4.94 0.04 |
| Energy | 3.85 0.05 | 3.00 0.03 | 3.09 0.03 |
| MB-MPO | MA ( ) | MA ( ) | MA ( ) | |
|---|---|---|---|---|
| Early stage of epoch: 10K | ||||
| Average total rewards | -1084 73 | 672 67 | -427 68 | -579 170 |
| Average feedback normalized | 38 3.95 | 13 0.35 | 36 5.91 | 26 1.91 |
| Average feedforward normalized | 7.11 0.06 | 7.33 0.04 | 5.99 0.05 | 6.86 0.06 |
| Average energy normalized | 3.63 0.017 | 3.20 0.016 | 3.27 0.15 | 3.50 0.010 |
| Middle stage of epoch: 100K | . |
| MB-MPO | Ours ( ) | Ours ( ) | Ours ( ) | |
|---|---|---|---|---|
| Early stage of epoch: 10K | ||||
| Average total rewards | -1245 801 | 99 180 | 90 102 | 100 139 |
| Average feedback normalized | 13.6 9.23 | 21.5 1.36 | 19.2 6.8 | 15.3 1.65 |
| Average feedforward normalized | 10.5 0.07 | 8.12 0.15 | 5.99 0.14 | 8.07 0.08 |
| Average energy normalized | 4.25 0.057 | 3.30 0.045 | 3.20 0.021 | 3.50 0.032 |
| Middle stage of epoch: 100K |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Training Steps | Mamba Baseline | + MA ( ) | + MA ( ) |
|---|---|---|---|
| 30M (Final) | 2453 110 | 2553 76 | 2498 123 |
| 10M (Mid) | 1985 156 | 2127 121 | 1965 136 |
| 2.5M (Early) | 467 288 | 759 231 | 415 245 |
| Task | Training Steps | DreamerV3 Baseline | + MA ( ) |
|---|---|---|---|
| HalfCheetah-run | 100K (Early) | 257 56 | 308 59 |
| 250K (Mid) | 408 86 | 394 81 | |
| 500K (Final) | 575 123 | 625 77 | |
| reacher-hard | 100K (Early) | 917 35 | 906 42 |
| 250K (Mid) | 874 20 | 930 17 | |
| 500K (Final) | 942 20 | 931 17 |
| MB-MPO | Ours ( ) | |
|---|---|---|
| Early stage of epoch: 10K | ||
| Average total rewards | 47 2 | 232 5.75 |
| Average feedback normalized | 168 5.26 | 290 1.82 |
| Average feedforward normalized | 0.45 0.026 | 0.11 0.003 |
| Average energy normalized | 0.26 0.0027 | 0.55 0.022 |
| Middle stage of epoch: 100K |
| MB-MPO | Ours ( ) | |
|---|---|---|
| Early stage of epoch: 10K | ||
| Average total rewards | 234 3.39 | 278 3.15 |
| Average feedback normalized | 99.7 2.15 | 137 5.25 |
| Average feedforward normalized | 0.11 0.003 | 0.03 0.001 |
| Average energy normalized | 1.52 0.058 | 1.14 0.015 |
| Middle stage of epoch: 100K |
| MB-MPO | Ours ( ) | |
|---|---|---|
| Early stage of epoch: 10K | ||
| Average total rewards | 325 226 | 615 45.76 |
| Average feedback normalized | 3 0.33 | 60.6 1.74 |
| Average feedforward normalized | 7.76 0.84 | 0.35 0.07 |
| Average energy normalized | 9.69 2.48 | 7.98 0.84 |
| Middle stage of epoch: 100K |
| MB-MPO | Ours ( ) | |
|---|---|---|
| Early stage of epoch: 10K | ||
| Average total rewards | 51 29 | 45 7.91 |
| Average feedback normalized | 4.62 0.21 | 3.1 0.7 |
| Average feedforward normalized | 6.47 0.25 | 3.05 0.19 |
| Average energy normalized | 5.00 0.42 | 3.94 0.29 |
| Middle stage of epoch: 100K |
| MB-MPO | Ours ( ) | |
|---|---|---|
| Early stage of epoch: 10K | ||
| Average total rewards | 297 10 | 245 10 |
| Average feedback normalized | 39.9 1.7 | 2.48 0.13 |
| Average feedforward normalized | 0.13 0.006 | 0.877 0.06 |
| Average energy normalized | 1.08 0.036 | 0.52 0.016 |
| Middle stage of epoch: 100K |
| MB-MPO | Ours ( ) | |
|---|---|---|
| Early stage of epoch: 10K | ||
| Average total rewards | 314 19 | 359 27 |
| Average feedback normalized | 59.1 5.24 | 15.9 1.06 |
| Average feedforward normalized | 0.285 0.003 | 0.136 0.002 |
| Average energy normalized | 1.01 0.082 | 0.053 0.0081 |
| Middle stage of epoch: 100K |