Event-Driven Proactive Robot Assistance through Vision-Language Reasoning
Organizations: The University of Osaka · Nanyang Technological University · The University of Tokyo
Abstract
Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we investigate an event-driven formulation of proactive assistance, where human--object interaction outcomes initiate assistive reasoning without user-provided task specifications at inference time. To this end, we propose an event-driven framework that monitors workspace state changes with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. A frozen pretrained Vision-Language Model (VLM) then uses its semantic priors to infer the task context, decide whether assistance is appropriate, and, when needed, generate a sequence of assistive actions from the observed transition. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs.We evaluate the same framework across three distinct real world tabletop collaboration tasks without task-specific training or fine-tuning. The event-driven framework achieves performance comparable to variants given user instructions.
Figures & tables
| Method | T1: Equation completion | T2: Tabletop sorting | T3: Tool introduction | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Appropriate | Succ. / QA / Fail | Latency | Appropriate | Succ. / QA / Fail | Latency | Appropriate | Succ. / QA / Fail | Latency | |
| Ours | 100.0 | 95.0 / 5.0 / 0.0 | 8.6 | 80.0 | 80.0 / 0.0 / 20.0 | 8.6 | 100.0 | 80.0 / 20.0 / 0.0 | 11.1 |
| w/o Pre-event Context | 93.3 | 91.7 / 1.7 / 6.7 | 6.8 | 60.0 | 53.3 / 6.7 / 40.0 | 7.4 | 90.0 | 80.0 / 10.0 / 10.0 | 9.4 |
| Explicit Request [ 14 , 18 ] | 98.3 | 91.7 / 6.7 / 1.7 | 7.2 | 86.7 | 80.0 / 6.7 / 13.3 | 7.2 | 100.0 | 53.3 / 46.7 / 0.0 | 5.0 |
| w/o Pre-event Context | 90.0 | 81.7 / 8.3 / 10.0 | 4.5 | 70.0 | 60.0 / 10.0 / 30.0 | 6.8 | 100.0 | 76.7 / 23.3 / 0.0 | 4.8 |
| Implicit Request [ 11 , 24 ] | 98.3 | 95.0 / 3.3 / 1.7 | 8.6 | 76.7 | 76.7 / 0.0 / 23.3 | 12.8 | 90.0 | 63.3 / 26.7 / 10.0 | 14.6 |
| Variant | T1 | T2 | T3 |
|---|---|---|---|
| Ours (Full) | 100.0 | 80.0 | 100.0 |
| w/o Event Triggering ( s) | 23.3 | 3.3 | 60.0 |
| w/o Event Triggering ( s) | 20.0 | 0.0 | 86.7 |
| w/o Early Stopping | 86.7 | 60.0 | 70.0 |
| Method | Optional Intervention | Required Abstention |
|---|---|---|
| Ours | 82.2 | 76.7 |
| Explicit Request | 83.3 | 50.0 |
| Implicit Request | 80.0 | 60.0 |