cs.ROSep 18, 2026

SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation

Authors: Saksham Singh, Zheyuan Hu, Max Sobol Mark, Jeffrey Yu, Zackory Erickson, Aviral Kumar

Organizations: Carnegie Mellon University

Abstract

Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement, but learning from sparse task-level rewards entails long credit-assignment horizons, difficult Bellman backups, and broad data-coverage requirements. We introduce SeeQ (Subtask-elicited Q-functions), which instead learns Q-values for the currently active subtask. This shortens the value-prediction horizon and enables effective learning with temporal-difference (TD) objectives. During training, subtask-level annotations present in offline robot data provide the decomposition and enable learning from broad, potentially suboptimal robot datasets. To eliminate the need for human annotations or modular subtask prediction systems at test time, our Q-function architecture is trained to autoregressively predict the active subtask in natural language before estimating its value. We instantiate SeeQ using a base vision-language backbone, pretrain it on diverse open-source robot manipulation data, and finetune it on downstream tasks. Across four real-world manipulation tasks on two bimanual robot platforms, the SeeQ value function substantially improves best-of-N policy steering.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. World Value Models for Robotic Manipulation

    Jun 23, 2026Zhihao Wang, Jianxiong Li, Yu Cui +4Efficient World-Action ModelRobotic Manipulation

  2. Freeform Preference Learning for Robotic Manipulation

    Jun 30, 2026Marcel Torne, Anubha Mahajan, Abhijnya Bhat +1Robot PoliciesRobotic Manipulation

  3. GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

    Jun 8, 2026Sriram Krishna, Ben Eisner, Haotian Zhan +5Robotic ManipulationRobot Systems