cs.AIOct 6, 2026

Sensor-Language-Action Models

Authors: Yuekai Xu, Zitao Shuai, Yuzhe Yang

Abstract

Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while remaining grounded in the underlying sensor evidence. We build a large-scale SLA benchmark consisting of datasets that span more than 116,000 individuals, 79 sensor modalities, and 60 action groups, together with a multi-faceted captioning pipeline that aligns user context, sensor dynamics, and action evidence. Building on this framework, we present OpenSLA, a unified SLA model for hierarchical action prediction, state understanding, and action explanation. Extensive experiments on real-world tasks in clinical prediction, operating rooms, and metabolic health verify its superior performance over the state-of-the-art. OpenSLA also demonstrates intriguing capabilities including language-guided evidence grounding and zero-shot generalization to unseen actions and cohorts.

Explore similar work

CardsList
  1. Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

    May 17, 2026Xinchen Jin, Aditya Chatterjee, Pranav Kumar +1VLM InterpretabilitySparse Autoencoders

  2. LARA: Latent Action Representation Alignment for Vision-Language-Action Models

    Jun 5, 2026Mengya Liu, Baoxiong Jia, Jiangyong Huang +2Latent Action LearningRobotic Manipulation