Robot State

Momentum

11 papers in the last four weeks, up 38% on the four weeks before. 0.1% of all new papers.

Jul 6Week of Sep 21

Latest papers 64

Feb 12, 2026cs.RO

LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation

This paper develops LongNav-R1, an end-to-end multi-turn reinforcement learning (RL) framework designed to optimize Visual-Language-Action (VLA) models for long-horizon navigation. Unlike existing single-turn paradigm, LongNav-R1 reformulates the navigation decision process as a continuous multi-turn conversation between the VLA policy and the embodied environment. This multi-turn RL framework offers two distinct advantages: i) it enables the agent to reason about the causal effects of historical interactions and sequential future outcomes; and ii) it allows the model to learn directly from online interactions, fostering diverse trajectory generation and avoiding the behavioral rigidity often imposed by human demonstrations. Furthermore, we introduce Horizon-Adaptive Policy Optimization. This mechanism explicitly accounts for varying horizon lengths during advantage estimation, facilitating accurate temporal credit assignment over extended sequences. Consequently, the agent develops diverse navigation behaviors and resists collapse during long-horizon tasks. Experiments on object navigation benchmarks validate the framework's efficacy: With 4,000 rollout trajectories, LongNav-R1 boosts the Qwen3-VL-2B success rate from 64.3% to 73.0%. These results demonstrate superior sample efficiency and significantly outperform state-of-the-art methods. The model's generalizability and robustness are further validated by its zero-shot performance in long-horizon real-world navigation settings. All source code is open-sourced at https://github.com/UMich-CURLY/LongNav-R1.
Feb 11, 2026cs.RO

H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model

World models are becoming central to robotic planning and control by predicting future state transitions. Existing approaches mainly rely on visual, latent, or language prediction, which can be difficult to ground in executable robot actions and prone to compounding errors over long horizons. In contrast, traditional robotic task and motion planning enables structured long-horizon reasoning through compact symbolic representations of world transitions, but typically lacks synchronized visual prediction. We propose Hierarchical World Model (H-WM), which jointly predicts logical and visual state transitions by combining a high-level logical world model with a low-level visual world model. The predicted logical actions and latent visual state transitions are jointly incorporated into Vision-Language-Action (VLA) models as intermediate state guidance for long-horizon task execution. Experiments on three long-horizon benchmarks and real robots show that H-WM consistently improves VLA's performance by stabilizing long-horizon execution and mitigating error accumulation. We also construct LIBERO-Logic, a frame-level aligned dataset that pairs visual observations and continuous robot states with logical actions and predicate-based logical states.
Feb 6, 2026cs.RO

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, preventing robot state from grounding instruction understanding or directing visual attention. We introduce ThinkProprio, which discretizes proprioception into VLM-vocabulary tokens and uses them jointly with the instruction to gate visual patches before VLM computation, steering the model toward action-relevant evidence while discarding redundant tokens early. We find that proprioception added as a passive conditioning signal leaves performance essentially unchanged; its value emerges when token-form state acts as an active query that, with the instruction, selects which visual patches the VLM processes. Systematic ablations show that VLM-vocabulary tokens outperform learned projectors as the state encoding, and that retaining only about \SI{12}{\percent} of the visual tokens surpasses on CALVIN ABC→\toD. Across CALVIN, LIBERO, and real-world manipulation, ThinkProprio reduces end-to-end inference latency while improving the matched full-token baseline.
Date pendingcs.RO

Retriever: Composing the Perception-Reasoning-Action Loop for Long-Horizon Manipulation

Building long-horizon robot agents requires composing closed-loop pipelines -- perception, belief update, planning, and control -- whose components run at different clocks and with variable latency. Today, these systems are often assembled with ad-hoc concurrency and pub/sub conventions that make timing and input-consumption semantics implicit, yielding schedule-dependent behavior that is hard to reproduce, debug, and reuse. Current solutions typically solve parts of this problem at either the algorithmic or the systems layer, but not both. In this work, we propose Retriever, which spans the entire stack: an asynchronous decision model, a programming model, a runtime, and an example closed-loop agent pipeline. Retriever represents an agent as a graph of stateful causal stream functions executed on explicit run clocks. We formalize this view via an asynchronous environment-agent loop over continuous-time streams and show that finite-memory causal policies can be represented by compositions of these operators. Retriever compiles these graphs into a runtime that supports multiple backends, enabling systematic debugging across running environments and deterministic replay from logged asynchronous data. We evaluate Retriever through a real-robot case study together with controlled studies of runtime overhead and deterministic replay behavior.