Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots
Organizations: School of Interactive Computing, Georgia Institute of Technology, USA
Abstract
Effective robot assistance beyond narrow roles and repetitive tasks requires robots to be proactive - to decide what needs to be done rather than waiting to be told. While proactivity is increasingly explored, it lacks a unified formulation, and work in the domain is typically evaluated offline against static human models that cannot capture the effect of a robot's actions on the environment and the user's own behavior. We introduce a unified formalism for proactive robot assistance, organize it into three levels, and provide a framework to address the highest level of unprompted proactive assistance. We then show that offline evaluation overstates performance in this setting, and contribute a closed-loop evaluation with a human model that adapts to the robot. Finally, we present a method, GAP, that instantiates our framework, learning from passive observation to anticipate user goals and act. Under closed-loop evaluation, prior state-of-the-art methods collapse, in some cases adding more work than they save, while GAP remains robust and substantially outperforms them.
Figures & tables
| Closed-Loop | |||
|---|---|---|---|
| Method | Offline | LLM User | Scripted User |
| SLaTe-PRO | 0.634 | 0.216 | 0.264 |
| STREAK | 0.677 | 0.002 | 0.296 |
| LLM User | Scripted User | |||||
|---|---|---|---|---|---|---|
| Method | Net Saved (%) | Perturbed (%) | F1 | Net Saved (%) | Perturbed (%) | F1 |
| SLaTe-PRO | 2.1 7.5 | 18.1 3.1 | 0.216 0.034 | 16.6 3.0 | 11.8 1.4 | 0.264 0.043 |
| STREAK | -159.9 50.3 | 10.6 10.8 | 0.002 0.001 | 20.3 7.3 | 48.3 19.0 | 0.296 0.105 |
| Plain LLM | -174.5 26.5 | 62.1 4.5 | 0.076 0.022 | 19.5 4.8 | 25.2 0.8 | 0.310 0.064 |
| GAP (ours) | 17.7 1.1 | 11.7 1.9 | 0.388 0.007 | 30.8 1.8 | 10.1 0.8 | 0.452 0.024 |
| Oracle | 44.4 1.6 | 8.2 1.1 | 0.617 0.015 | 51.3 0.0 | 8.8 1.9 | 0.651 0.006 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Distribution Shift at Test time only | Distribution Shift during Training (main results) | |||||
|---|---|---|---|---|---|---|
| Method | Net Saved (%) | Perturbed (%) | F1 | Net Saved (%) | Perturbed (%) | F1 |
| SLaTe-PRO | 24.2 5.4 | 10.8 3.0 | 0.345 0.071 | 16.6 3.0 | 11.8 1.4 | 0.264 0.043 |
| STREAK | 0.4 0.6 | 66.3 57.4 | 0.009 0.015 | 20.3 7.3 | 48.3 19.0 | 0.296 0.105 |
| GAP (ours) | 36.1 1.8 | 16.0 0.7 | 0.496 0.022 | 30.8 1.8 | 10.1 0.8 | 0.452 0.024 |