Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement, but learning from sparse task-level rewards entails long credit-assignment horizons, difficult Bellman backups, and broad data-coverage requirements. We introduce SeeQ (Subtask-elicited Q-functions), which instead learns Q-values for the currently active subtask. This shortens the value-prediction horizon and enables effective learning with temporal-difference (TD) objectives. During training, subtask-level annotations present in offline robot data provide the decomposition and enable learning from broad, potentially suboptimal robot datasets. To eliminate the need for human annotations or modular subtask prediction systems at test time, our Q-function architecture is trained to autoregressively predict the active subtask in natural language before estimating its value. We instantiate SeeQ using a base vision-language backbone, pretrain it on diverse open-source robot manipulation data, and finetune it on downstream tasks. Across four real-world manipulation tasks on two bimanual robot platforms, the SeeQ value function substantially improves best-of-N policy steering.
Figures & tables
Figure 2 : Overview of the SeeQ value-function architecture. The model first rolls out the active subtask at state s in text using a next-token prediction loss, and then predicts a scalar value Qθ(s,a,l,l~) for the input state-action pair conditioned on the task instruction and the subtask. Ground-truth subtask annotations supervise the rollout during training; at inference the model conditions on its own predicted subtask.
Figure 3 : Our bimanual real-robot evaluation tasks. Top-camera snapshots from dataset demonstrations of the four real-world tasks. Three tasks are set up on a bimanual xArm-7 platform, and one uses a bimanual YAM platform. See Appendix A.1.5 for task definitions and evaluation protocols.
Figure 4 : Qualitative value predictions for our approach compared to basline approaches. Panels (b) and (c) both show the values of selected best-of- N action chunks in a SeeQ evaluation. (a) All critics evaluate recorded dataset action chunks. On a held-out RoboCOIN trajectory, task-level TD provides almost no signal for most of the trajectory, consistent with compounding bootstrapping errors, while MC predictions are noisy, potentially reflecting spurious correlations. (b) lid-sealing: The value drops at frame 3 after a failed flap closure, followed by recovery at frame 4. (c) grocery-packing: Value drops at frames 1 and 3 correspond to failed grasps, followed by recovery in both cases; BC often struggles repeatedly with objects such as the Pringles bag. The final subtask switch appears delayed because subtask decoding runs less frequently than action replanning. Solid curves denote SeeQ and TD-BoN; dashed curves denote SARSA and MC. Red vertical dashed lines mark SeeQ -predicted subtask switches and are omitted from the task-level value plot.
Task
Data
Base
SeeQ
shirt-hang
RaC
10/24
22/24
lid-sealing
RaC
10/24
15/24
grocery-packing
Teleop
9/24
17/24
LEGO-disassembly
Teleop
5/24
10/24
Table 1 : Policy steering results. Success rates over 24 trials for the base policy and for the same policy steered by SeeQ with best-of- N action selection ( N=8 ).
Value function
Robot pretraining
VLM initialization
Success rate
Base policy
—
—
10/24
Task-specific (ResNet-50 + MLP, TD-BoN)
✗
✗
16/24
SeeQ minus pretraining
✗
✓
8/24
SeeQ (Ours)
✓
✓
22/24
Table 3 : Importance of pretraining and VLM initializations for SeeQ , evaluated on shirt-hang. Columns indicate whether the value function uses robot data pretraining and VLM initialization. Observe that both robot data pretraining and base VLM initialization are important for the success of SeeQ .
Variant
Success rate
Base policy
10/24
Subtask-level TD No subtask prediction, no subtask conditioning
8/24
Subtask-level TD Subtask prediction, no subtask conditioning
15/24
SeeQ (Ours)
22/24
Table 4 : Ablation of subtask elicitation on shirt-hang. Subtask prediction supervision improves success, with further gains from explicitly conditioning value estimation on the predicted subtask.
Figure 5 : Action and image sensitivity on six complete held-out shirt-hang trajectories. (a) Action-gradient norms. (b) Joint image-gradient norms across three cameras. All three critics use the same 12,582 frames and eight cached policy candidates; gradients are evaluated at each critic’s selected candidate. No pretraining denotes SeeQ finetuned directly from PaliGemma. Bar heights are percentages of frames in shared logarithmically spaced bins; dashed lines denote medians.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Q-values relative to recorded returns during RoboCOIN pretraining. Gray: batch-mean Q−MC ; red: EMA with a 1,000 -step half-life; shading: ±1 standard deviation of the exponentially weighted prediction-minus-return differences. (a) SeeQ with subtask returns shows a near-zero gap after initial underestimation. (b) Task-level TD-BoN persistently underestimates task returns.
Figure 7 : Subtask conditioning aligns value resets with predicted subtask changes. Both critics evaluate recorded dataset action chunks. Top: predicted values along a held-out shirt-hang trajectory. Bottom: enlarged views of the five shaded regions; colored dashed lines mark each critic’s predicted switch nearest the panel center. Fall span counts frames between the first 5% and 95% of the local peak-to-trough drop, using the maximum value in the 20 frames before that critic’s switch and the minimum from the switch through 20 frames after it. A drop crossing both thresholds in one frame has span 0 . Panel titles report SeeQ versus no conditioning.
Figure 8 : Comparison of the value functions in Section 4.2 . All critics evaluate recorded dataset action chunks. (a) In this held-out episode, the Nesquik container must be reoriented in frame 2 because the soda can obstructs it, causing SeeQ ’s value to drop. In the final frame, MC predictions are noisy despite smooth progress. (b) In this held-out episode, the value jumps in frame 1 as the gripper reorients to grasp the hanger correctly. In frame 4, the left gripper unintentionally releases the collar, causing the value to drop. Frame 5 again shows noisy MC predictions.
Figure 9 : Comparison of the value functions in Section 4.3 on a held-out episode. All critics evaluate recorded dataset action chunks. In frame 3 (red), the right gripper releases the collar, causing the value to drop. The ResNet-50 value peaks about 0.7 s before this release, so the evaluated action chunk includes the failure; this indicates an incorrect value estimate, unlike the VLM critics in this case. SeeQ generally produces smoother predictions than SeeQ without robot pretraining and the task-specific ResNet-50 critic, indicating better generalization.
Task
Training steps
LR decay steps
Batch size
shirt-hang
60k
200k
128
lid-sealing
40k
60k
256
grocery-packing
70k
70k
256
LEGO-disassembly
20k
20k
256
Appendix
Table 5 : Downstream base-policy hyperparameters. Training steps refer to the policy checkpoint used in evaluation; LR decay length refers to the full configured cosine schedule.
Critic
Return scope
γ
NTP weight
Subtask conditioning
SeeQ (TD-BoN)
Subtask
0.999
0.1
Yes
SARSA
Subtask
0.999
0.1
Yes
MC
Task
0.9995
0
No
TD-BoN
Task
0.9995
0
No
NTP, no conditioning
Subtask
0.999
0.1
No
Neither NTP nor conditioning
Subtask
0.999
0
No
Appendix
Table 6 : Critic pretraining variants. All rows use 230 k pretraining steps and the shared schedule above. Conditioning refers to the critic’s value prediction, independently of the base-policy prompt.
Critic
Downstream task
Finetuning steps
Checkpoint step
SeeQ
shirt-hang , lid-sealing , grocery-packing
20k
250k
SeeQ
LEGO-disassembly
10k
240k
Task-level MC
shirt-hang , grocery-packing
20k
250k
Task-level TD-BoN
shirt-hang , grocery-packing
20k
250k
Subtask-level SARSA
shirt-hang , grocery-packing
20k
250k
NTP, no conditioning
shirt-hang
10k
240k
Appendix
Table 7 : Downstream critic checkpoints. All rows initialize from the corresponding 230 k pretrained critic and use a 20 k LR decay schedule.
Figure 10 : Histogram of RoboCOIN subtask durations in seconds.
Agilex Cobot Magic (63 tasks)
box-storage-chopsticks
cap-the-pen-a
catch-the-ball
classification-of-fruits-and-vegetables
classification-of-fruits-and-vegetables-a
classification-of-tableware
clean-blackboard
clean-up-the-tableware
clear-the-desktop
close-book
close-button
cube-reset
cut-banana
desktop-organization
drawer-storage-mineral-water
fold-clothes
fold-the-towel
fold-towel-a
Appendix
Table 8: The 131 RoboCOIN task datasets used for pretraining. Task identifiers are grouped by embodiment, with the embodiment prefix omitted and underscores rendered as hyphens. Released spellings and variant suffixes are retained.
Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. Mathematically, accurate value estimation demands deep temporal understanding, requiring models to both ground the current belief using historical context and plan over future outcomes. However, most existing robotic value models are built on Vision-Language Model (VLM) backbones that are pretrained primarily on static or temporally sparse visual observations, lacking the requisite temporal modeling capabilities for value estimation. Unlike VLMs, world models naturally excel at temporal modeling and future planning, making them ideal foundations for learning generalizable value functions. Driven by this insight, we marry world models with value estimation to construct a new generalist robotic value model, World Value Model (WVM), that offers accurate task progressions to assess data quality. On standard benchmarks, WVM delivers state-of-the-art (SOTA) Value-Order Correlation (VOC) results. Complementing standard evaluation suites that contains only expert data, we further introduce Suboptimal-Value-Bench, a multi-embodiment benchmark consisting of 800 suboptimal trajectories with high-fidelity, human-labeled frame annotations. Our evaluations show that WVM maintains its SOTA performance on Suboptimal-Value-Bench, establishing its robustness in handling both expert and suboptimal data. When deployed for policy learning, WVM improves manipulation performance across various policy extraction approaches in both simulated and real-world deployment, providing robust guidance for learning from mixed-quality data.
Zhihao Wang, Jianxiong Li, Yu Cui +4
ByteDance Seed · Tsinghua University · Peking University
Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal. We introduce Freeform Preference Learning (FPL), a method for learning robot policies from freeform human preferences. Rather than asking annotators which of two trajectories is better overall, FPL lets them define natural-language preference axes, such as speed, safety, quality of placement, or carefulness, and provide pairwise preferences along each axis. These annotations are used to learn a language-conditioned reward model that maps a trajectory and preference label to an axis-specific reward. We use this model to train a reward-conditioned policy that optimizes across the multiple human-specified dimensions. Across four real-world and two simulated long-horizon manipulation tasks, FPL improves over sparse-reward and binary-preference methods by 38 percentage points. Beyond improved performance, FPL learns dense progress signals without explicit subtask segmentation, shows compositionality of behavior not present in the data, and allows users to steer the policy towards different behaviors at test time without retraining. Blog post with videos available at https://freeform-pl.github.io/fpl.website/
We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii) a low-level goal-conditioned controller that executes embodiment-specific actions. To condition image-based policies on 3D goals, we introduce a simple spatial interface that projects predicted goals into the image plane and represents them as end-effector heatmaps. Across a suite of manipulation tasks, this hierarchical factorization consistently improves performance and robustness compared to a flat Diffusion Policy. Further, we show that this hierarchical interface also makes it easy to incorporate human demonstrations without relying on (noisy) action retargeting. As sub-goals are largely embodiment-agnostic, we train the high-level policy on human video to specify how learned skills should be applied and composed, while keeping the low-level policy trained purely on robot data. This hierarchy enables adaptation to novel objects and task variations using a small number of human demonstrations.
Sriram Krishna, Ben Eisner, Haotian Zhan +5
Robotics Institute, Carnegie Mellon University · UMass Amherst