cs.ROOct 6, 2026

Event-Driven Proactive Robot Assistance through Vision-Language Reasoning

Authors: Fengkai Liu, Hao Su, Haozhuang Chi, Rui Geng, Congzhi Ren, Xuqing Liu, Chenfei Xu, Yuichi Ohsita, +1 more

Organizations: The University of Osaka · Nanyang Technological University · The University of Tokyo

Abstract

Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we investigate an event-driven formulation of proactive assistance, where human--object interaction outcomes initiate assistive reasoning without user-provided task specifications at inference time. To this end, we propose an event-driven framework that monitors workspace state changes with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. A frozen pretrained Vision-Language Model (VLM) then uses its semantic priors to infer the task context, decide whether assistance is appropriate, and, when needed, generate a sequence of assistive actions from the observed transition. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs.We evaluate the same framework across three distinct real world tabletop collaboration tasks without task-specific training or fine-tuning. The event-driven framework achieves performance comparable to variants given user instructions.

Figures & tables

Explore similar work

CardsList
  1. ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

    Jun 16, 2026Tianyi Lu, Hui Zhang, Zijie Diao +8Diffusion-Based Vision-Language-ActionsRobotic Manipulation

  2. A Conversational Framework for Human-Robot Collaborative Manipulation with Distributed Generative AI models

    Jun 4, 2026Arash Ghasemzadeh Kakroudi, Roel PietersHuman-Robot CollaborationRobotic Manipulation

  3. Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical Dual-Correction

    Feb 16, 2026Taichi Kato, Takuya Kiyokawa, Namiko Saito +1Human-Robot CollaborationRobotic Manipulation