Behavioral Foundation Models (BFMs) aim to solve a wide range of downstream tasks without test-time policy learning by inferring a task vector from the reward function. While efficient, the retrieved policies are often suboptimal because of how this task vector is inferred, typically with ordinary least squares (OLS). OLS minimizes reward reconstruction error but leaves the ordering of rewards unconstrained, which can bias the successor measure of the retrieved zero-shot policy away from that of the optimal policy. In this work, we propose BLS, an efficient test-time inference method that balances minimizing reward reconstruction error with reducing successor-measure mismatch. Theoretically, we provide a suboptimality gap upper bound characterized by both successor-measure and reward-function residuals. Empirically, we evaluate BLS on top of state-of-the-art BFMs across benchmarks for locomotion, manipulation, and humanoid control. BLS outperforms existing task inference baselines with negligible computational overhead. Project page: https://embodiedai-ntu.github.io/BLS
Figures & tables
Figure 1: Ordinary least squares (OLS) aims to minimize the total errors of reconstructed rewards. Its performance highly depends on the quality of state representations, the state coverage within the dataset, and the distribution of rewards. Because OLS can be biased toward the majority of states, low-reward states ( N ) may be assigned with high-value rewards and vice versa for high-reward states ( G ). The successor measure of the retrieved zero-shot policy can therefore diverge from that of the optimal policy. BLS instead adopts an objective that reduces the reward reconstruction errors while enlarging the margin between high- and low-reward states. Thus, the successor measure of the retrieved zero-shot policy is encouraged to stay closer to the optimal policy’s.
Algorithm 1 BLS
Laplacian
ICVF
HILP
FB
TD-JEPA
OLS
BLS
OLS
BLS
OLS
BLS
OLS
BLS
OLS
BLS
antmaze-mn
48.1 ± 2.9
67.3 ± 8.1
66.3 ± 1.2
69.2 ± 6.7
74.3 ± 1.4
87.7 ± 1.5
77.5 ± 2.6
88.8 ± 4.2
63.1 ± 0.8
83.2 ± 7.7
antmaze-ln
7.6 ± 1.8
10.9 ± 2.9
56.0 ± 4.1
65.5 ± 4.9
49.2 ± 2.1
63.7 ± 1.8
52.4 ± 4.0
63.5 ± 9.2
50.7 ± 2.4
64.3 ± 7.0
antmaze-ms
15.1 ± 7.8
34.7 ± 3.4
28.8 ± 5.6
36.5 ± 9.1
26.3 ± 9.1
31.3 ± 2.3
69.9 ± 1.7
79.6 ± 2.1
54.5 ± 4.8
59.5 ± 2.4
antmaze-ls
15.2 ± 3.9
15.3 ± 0.6
5.5 ± 0.8
4.0 ± 0.7
11.9 ± 1.2
13.5 ± 0.8
51.5 ± 2.9
45.7 ± 4.6
39.2 ± 5.9
45.1 ± 5.9
antmaze-me
3.7 ± 3.3
7.6 ± 1.7
0.3 ± 0.5
0.0 ± 0.0
2.4 ± 2.6
6.7 ± 1.5
50.3 ± 2.0
49.2 ± 4.9
15.1 ± 5.5
29.2 ± 7.0
Table 1: Zero-shot success rates (%) on OGBench across five feature representations under proprioceptive and pixel-based observations. BLS improves the average performance of all representations.
Table 4
Figure 2: Comparison of OLS, BLS, and GCRL on OGBench using HILP. The hatched segments indicate the improvement of BLS over OLS. BLS achieves performance comparable to GCRL on average, while their relative performance varies across tasks.
Figure 3: Dense-reward results on DMC (top) and HumEnv (bottom). Solid bars show the return of the OLS policy, and hatched segments indicate the change after applying BLS. DMC reports average return across tasks for each domain and feature representation, while HumEnv reports expert-normalized return on 45 motion-tracking tasks, with FB-CPR feature representations.
Figure 4: Sensitivity to λ with the HILP encoder: sparse rewards (left two) prefer a small λ , and dense rewards (right two) prefer a large one.
Figure 5: Sensitivity to the reward threshold for defining G on DMC with TD-JEPA. We report the paired mean return gain of BLS over OLS across 20 tasks as the threshold varies.
d
25
50
100
200
500
1000
OLS
70.9 ± 2.1
74.3 ± 1.4
59.2 ± 1.7
25.3 ± 1.5
9.2 ± 0.4
22.0 ± 5.7
BLS
84.7 ± 2.0
87.7 ± 1.5
87.5 ± 2.0
82.1 ± 9.0
79.6 ± 8.4
82.7 ± 16.7
Table 4: Success rate as the feature dimension d varies.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Laplacian
Task
OLS
BLS
ReLA
LoLA
antmaze-mn
48.1 ± 2.9
67.3 ± 8.1
39.6 ± 8.0
48.5 ± 9.2
antmaze-ln
7.6 ± 1.8
10.9 ± 2.9
7.3 ± 6.1
5.6 ± 1.1
antmaze-ms
15.1 ± 7.8
34.7 ± 3.4
12.9 ± 8.5
10.8 ± 4.2
antmaze-ls
15.2 ± 3.9
15.3 ± 0.6
4.9 ± 5.8
12.9 ± 5.7
antmaze-me
3.7 ± 3.3
7.6 ± 1.7
0.0 ± 0.0
3.5 ± 3.6
Appendix
Table 5: OGBench success rates (%) across four feature representations.
Task
OLS
ZOL
BLS@OLS
BLS@ZOL
ReLA
LoLA
antmaze-mn
77.5 ± 2.6
84.5 ± 1.6
88.8 ± 4.2
84.4 ± 4.5
70.9 ± 3.0
80.5 ± 2.4
antmaze-ln
52.4 ± 4.0
41.5 ± 0.8
63.5 ± 9.2
68.7 ± 2.4
39.7 ± 3.1
52.8 ± 5.8
antmaze-ms
69.9 ± 1.7
51.9 ± 1.2
79.6 ± 2.1
68.1 ± 3.6
76.1 ± 1.9
70.0 ± 2.2
antmaze-ls
51.5 ± 2.9
37.9 ± 5.1
45.7 ± 4.6
36.0 ± 4.6
29.2 ± 1.1
53.7 ± 3.8
antmaze-me
50.3 ± 2.0
49.7 ± 3.1
49.2 ± 4.9
48.8 ± 1.4
45.5 ± 0.5
47.6 ± 7.3
cube-single
50.4 ± 0.8
46.4 ± 6.8
54.0 ± 1.8
43.7 ± 6.5
54.5 ± 2.9
54.5 ± 3.6
Appendix
Table 6: OGBench success rates (%) with the FB representation, including the FB-specific ZOL baseline and the online adaptation baselines ReLA and LoLA.
Task
ICVF
HILP
FB
TD-JEPA
Avg.
walker-stand
988 / 961
954 / 959
938 / 944
956 / 966
959.0 / 957.5
walker-walk
850 / 884
943 / 932
921 / 919
754 / 834
867.0 / 892.2
walker-run
307 / 279
332 / 346
353 / 333
316 / 337
327.0 / 323.8
walker-spin
960 / 972
982 / 968
978 / 977
954 / 962
968.5 / 969.8
cheetah-walk
979 / 979
255 / 651
876 / 974
967 / 983
769.2 / 896.8
cheetah-walk-backward
986 / 968
874 / 969
985 / 963
981 / 982
956.5 / 970.5
Appendix
Table 7: Task-level DMC episode return under the trust-region configuration used in the main experiments. Each cell reports OLS / BLS, averaged over three random seeds. The final column averages over the four feature representations used in Figure 3 .
OLS
ΨG
WR
B(g)
BLS@ B(g)
success ↑
0.144 ± 0.020
0.164 ± 0.008
0.148 ± 0.032
0.500 ± 0.022
0.516 ± 0.032
proximity ↑
0.455 ± 0.023
0.460 ± 0.022
0.453 ± 0.022
0.681 ± 0.009
0.699 ± 0.009
distance ↓
3.559 ± 0.116
3.309 ± 0.116
3.533 ± 0.094
2.668 ± 0.029
2.630 ± 0.038
Appendix
Table 8: Goal-reaching on HumEnv ( 50 target poses, FB-CPR).
Checkpoint
WR
ΨG
OLS
BLS@OLS
BLS@WR
S-1
0.677
0.673
0.509
0.686
0.744
S-2
0.575
0.573
0.491
0.638
0.696
S-3
0.575
0.560
0.480
0.641
0.696
S-4
0.624
0.622
0.524
0.655
0.679
S-5
0.615
0.640
0.477
0.622
0.686
Mean
0.613 ± 0.038
0.614 ± 0.042
0.496 ± 0.018
0.648 ± 0.022
0.700 ± 0.023
Appendix
Table 9: Expert-normalized return on HumEnv across five official FB-CPR checkpoints.