Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-reality demonstrations per task with synchronised multi-camera views and full-body execution records. We benchmark vision-language-action fine-tuning, imitation learning, demo-driven reinforcement learning, and cold-start coding agents given the interaction budget of online reinforcement learning. With the same onboard views, proprioception and whole-body controller for every method, vision-language-action fine-tuning has the highest nine-task mean, and agent-developed programs outperform every demo-driven reinforcement learning baseline on this mean and lead on bimanual reaching. Cross-workspace stacking remains open, π0.5 stays low on pick-box, and multi-object transport is hard for imitation learning, demo-driven reinforcement learning and coding agents. All environments, human demonstrations, and evaluation traces are open-sourced at https://github.com/swirl-uk/BiGym2.
Figures & tables
Fig. 4 : Body posture during cutlery and plate loading. Left: BiGym; right: BiGym 2.0. Each panel overlays two states from one demonstration, with the earlier pose faint and the later pose opaque. Both systems adjust working height. G1 also pitches its torso forward during plate loading. Robots and task layouts follow their respective benchmarks.
VLA
Imitation learning
Demo-driven RL
Coding agent (strict)
Task
π0.5
ACT
DP
CQN-AS
DrQ-v2+
DEAS
Opus 5.5
Astra
Pick box
0 19
0 61 ± 4
0 10 ± 3
0 48 ± 2
00 0 ± 0
0 14 ± 3
0 71 ± 16
0 54 ± 22
Reach multi-modal
0 90
0 81 ± 1
0 84 ± 1
0 19 ± 16
00 0 ± 0
0 79 ± 2
100 ± 0
0 89 ± 7
Reach dual
0 58
0 30 ± 0
0 27 ± 0
00 3 ± 1
00 0 ± 0
0 34 ± 1
0 78 ± 10
0 89 ± 7
Drawer close
100
100 ± 0
100 ± 0
100 ± 0
0 99 ± 0
100 ± 0
100 ± 0
0 99 ± 1
Drawer open
0 99
0 92 ± 0
0 98 ± 0
0 89 ± 2
00 0 ± 0
0 79 ± 4
0 99 ± 1
0 99 ± 1
TABLE II : Success rates (%) on nine representative tasks. ACT, DP, CQN-AS, DrQ-v2+ and DEAS: mean ± standard error (SE) across three training runs, each averaging its last five checkpoints with 100 episodes each. π0.5 : one training run per task, pooled over its last five checkpoints with 50 episodes each. Coding agents: three independent development sessions each of Claude Opus 5.5 and GPT-6 Astra under the strict interface ( 84×84 onboard views, direct joint actions, no kinematics tools), 100 episodes per frozen program.
Task
ACT
CQN-AS
Reach single
0 89 ± 0
0 84 ± 1
Flip cup
0 62 ± 1
0 42 ± 9
Flip cutlery
0 46 ± 1
0 55 ± 2
Stack blocks
00 0 ± 0
00 1 ± 0
Load cutlery
0 11 ± 1
0 10 ± 3
Load plates
0 60 ± 1
0 34 ± 9
TABLE III: ACT and CQN-AS success rates (%) on the remaining eleven tasks. Mean ± SE across three training runs, each averaging its last five checkpoints with 100 episodes each.
Success rate (%)
Cost/session ($)
Task
Strict
Tools
Δ
Strict
Tools
Reach multi
100 ±
0
100 ±
0
0
0 3.2
0 2.1
Reach dual
0 94 ±
2
0 99 ±
1
5
0 5.3
0 5.2
Drawer close
100 ±
0
100 ±
0
0
0 1.8
0 2.3
Drawer open
0 55 ±
27
0 99 ±
1
44
10.5
0 5.2
Move plate
0 39 ±
5
0 41 ±
26
2
10.8
11.7
TABLE IV: Coding-agent (GPT-6 Astra) performance and development cost. Strict vs. tool-augmented interfaces across nine tasks, from a separate set of sessions on a pre-release build of the environment, so strict scores differ from Table II . Task costs report mean per-session expenditure, with 27-session totals below.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Act. dim
Horizon
Success predicate (held for 1 s)
Reaching (3 tasks)
reach_target_single †
20
18 s (900)
Left end-effector (pinch-centre) site within 5 cm of the target.
reach_target_multi_modal
20
14 s (700)
Either end-effector site within 5 cm of the target.
reach_target_dual
20
14 s (700)
Each end-effector site within 5 cm of its own target, simultaneously.
Tabletop (5 tasks)
move_plate
20
34 s (1700)
Plate within 5 cm of a target-rack slot, normal within 20∘ of the slot axis, touching the rack and not the table, released.
Appendix
TABLE A1: Episode budgets and success predicates of the 20 tasks. Horizon is the episode budget in seconds (control steps at 50 Hz). Every predicate must hold for 1 s. Normalised joint travel is 0 at one end of a joint’s range and 1 at the other. Tasks marked † are outside the nine-task main comparison.
Fig. A1 : The 20 tasks of BiGym 2.0. MuJoCo renders of one demonstration per task: the state after reset (left) and the final frame (right). The coloured rule above each pair marks the scene family, as in Fig. . The final frame marks the success criterion: the 5 cm tolerance sphere for reaching, the object’s recorded displacement for placement tasks and the measured joint extent against its threshold for articulated fixtures. Table A1 gives the full predicates.
Fig. A2 : Onboard cameras. (a) MuJoCo renders of the G1 at the start of a plate-transport demonstration; the green markers are the three cameras. The head camera sits on torso_link and each wrist camera on its wrist_yaw_link , 4.9 cm from the pinch centre between the finger pads. All three have a 60∘ field of view, horizontally and vertically, because the image is square. (b) Stored 84×84 observations from one demonstration of each of three tasks, taken when a finger pad first touches an object (for reaching, at success), upscaled 6× by nearest neighbour without other processing.
Channel
Slice
Content
Physical range
Policy range
Action (20-D; 21-D with torso pitch ∗ ), 50 Hz
Forward velocity vx
[0]
Body command to GR00T-WBC
[−0.35,0.35] m/s
[−1,1]
Lateral velocity vy
[1]
Body command to GR00T-WBC
[−0.25,0.25] m/s
[−1,1]
Base height zbase
[2]
Body command to GR00T-WBC
[0.40,1.00] m
[−1,1]
Yaw rate ωz
[3]
Body command to GR00T-WBC
[−0.50,0.50] rad/s
[−1,1]
Torso pitch θpitch ∗
[4]∗
Waist pitch reference
[−0.20,0.80] rad
[−1,1]
Appendix
TABLE A2: Action and observation layout of the G1. The standard configuration has 20 action and 50 proprioceptive dimensions; torso-pitch tasks ( ∗ ) have 21 and 56. Command ranges are those recorded in the demonstration metadata.
Control steps/s
Reset
Task
No render
3×842
3×2242
(s)
Move plate
399 / 360
223 / 210
211 / 196
0.47
Reach dual
528 / 550
278 / 268
256 / 270
0.36
Dishwasher close
457 / 405
232 / 210
225 / 204
0.43
Appendix
TABLE A3 : Single-environment throughput in control steps per second (mean ± std over three repeats of 1,000 steps; hold action / demonstration replay) and reset time (mean ± std, 30 resets). Measured on an idle RTX 5090, with the process pinned to one core of a Threadripper PRO 7975WX.
Fig. A3 : Kinematic distributions of the 1,200 human VR demonstrations : 20 tasks × 60 successful episodes. A : episode duration (median 14.4 s). B : planar base travel, from 0.06 m near-stationary manipulation to 8.10 m cross-workspace transport. C : peak absolute pelvis lean within an episode, not a net or mean lean. D : physical pelvis height modulation, which is the executed travel rather than the command; the dashed line marks the 1.3 cm gait floor of the 11 tasks whose height command never moves. Boxes give the median and interquartile range, whiskers extend to 1.5× IQR and dots are episodes beyond them. Tasks are ordered by median base travel, so one row reads across all four panels; A, B and D are log-scaled. † Plate loading is the only task whose commanded torso-pitch action ever leaves 0 rad (58/60 episodes, peak 0.8 rad), yet every episode of all 20 tasks leans 6.4 – 27.9∘ .
Fig. A4 : Replay determinism. One move_plate demonstration (seed 18) replayed open-loop 20 times from its recorded engage snapshot and 20 times after a seed-only reset. Each curve is the largest difference to the first replay of the same condition, over the 29 actuated joint angles (left) and the plate position (right). From the snapshot, the replays are identical at every step. After a seed-only reset, the joint angles start 10−6 rad apart and the difference grows along the trajectory; the plate positions coincide until the gripper reaches the plate at about 4 s and end at most 0.73 mm apart. All 40 replays succeed.
Hyperparameter
ACT
DP
CQN-AS
DrQ-v2+
DEAS
π0.5
Policy representation
Transformer CVAE
1D temporal U-Net (FiLM)
C2F C51 critic, GRU over chunk
Actor + distributional twin critic
Distributional critics + AWR actor
PaliGemma VLM + flow-matching expert
Visual encoder
ResNet-18 (ImageNet, frozen BN)
ResNet-18 (ImageNet, frozen BN), spatial softmax
4-layer CNN per camera
4-layer CNN per camera
4-layer CNN per camera
SigLIP-So400M
Proprioception input
Linear token
Projected, concatenated
Projected, concatenated
Projected, concatenated
Projected, concatenated
256-bin discretised text tokens
Action chunk ( K )
32
16
32
1
16
50
Executed per prediction ( H )
1
8
1
1
16
16
Temporal ensemble
Yes ( m=0.01 )
No
Yes ( m=0.01 )
No
No
No
Appendix
TABLE A4: Hyperparameters of the evaluated methods. Non-VLA values are resolved from the run configurations of the main-table experiments. The online methods (CQN-AS, DrQ-v2+) sample 256 replay and 256 demonstration transitions per update and perform one update per environment step.
Setting
Value
Base checkpoint
Released π0.5 ( lerobot/pi05_base )
Adaptation
Full supervised fine-tuning (no LoRA)
Demonstrations
60 successful VR episodes per task
Cameras
Head, left wrist, right wrist; 84×84 RGB (model 224×224 )
Proprioception
50-D (no torso pitch) or 56-D (pitch tasks)
Action
Relative arm delta; base and gripper absolute
Appendix
TABLE A5: π0.5 fine-tuning settings. One training run per task.
Task
Instruction
Pick box
Pick up the box from the side table and place it on the counter.
Reach multi-modal
Reach the target with either wrist.
Reach dual
Reach the two targets, one with each wrist.
Drawer close
Close the top drawer of the kitchen cabinet.
Drawer open
Open the top drawer of the kitchen cabinet.
Move plate
Move the plate between two draining racks.
Appendix
TABLE A6: Language instructions used during π0.5 fine-tuning and evaluation. † Not in the nine-task main comparison.
ACT
DP
CQN-AS
DEAS
Task
Peak
Last-5
Δ
Peak
Last-5
Δ
Peak
Last-5
Δ
Peak
Last-5
Δ
Pick box
71.7
61.3
+10.4
18.3
9.9
+8.4
55.0
47.6
+7.4
26.7
13.7
+13.0
Reach multi-modal
83.3
80.6
+2.7
87.0
83.7
+3.3
28.3
18.9
+9.5
85.7
79.3
+6.4
Reach dual
33.0
29.9
+3.1
30.7
26.8
+3.9
10.7
2.9
+7.8
40.0
33.8
+6.2
Drawer close
100.0
99.8
+0.2
100.0
100.0
+0.0
100.0
99.9
+0.1
100.0
100.0
+0.0
Drawer open
94.3
92.3
+2.1
99.3
98.1
+1.2
93.3
89.1
+4.2
90.7
78.7
+12.0
Appendix
TABLE A7: Peak checkpoint versus last-five mean on the nine main tasks. Peak: mean over three runs of each run’s best checkpoint (100 episodes per checkpoint). Last-5: the main-table score. Δ=Peak−Last-5 in percentage points. DrQ-v2+ is omitted (99% last-five mean on drawer closing, 0% on the other eight tasks).
Strict Interface (84 × 84 RGB, direct joint targets)
TABLE A8: Per-session success and cost of GPT-6 Astra under the strict and tool-augmented interfaces , from the same sessions as Table IV . Each session’s frozen program is evaluated on 100 hidden seeds; Mean ± SE is across the three sessions. Cost is the mean per-session API-equivalent cost. On block stacking, ACT reaches 0.13% and tool-augmented Astra 6.3% (three-session mean; best session 19%).
Task
API-equivalent cost ($)
Input tokens
Run 1
Run 2
Run 3
median (M)
Reach multi-modal
2.72
4.33
2.65
1.5
Reach dual
3.09
6.72
6.08
3.7
Move plate
11.19
11.36
9.71
8.4
Move two plates
22.87
31.23
21.78
20.7
Dishwasher close
13.22
21.63
24.17
16.6
Appendix
TABLE A9: Per-session token use and cost of GPT-6 Astra under the strict interface on the pre-release build, from the same sessions as Table IV . Three sessions per task.
Claude Opus 5.5
GPT-6 Astra
Task
s1
s2
s3
s1
s2
s3
Pick box
44 ∗
99 ∗
71 ∗
64
85
12
Reach multi-modal
100 ∗
100 ∗
100 ∗
75
97
95
Reach dual
71 ∗
99 ∗
65 ∗
100
93
75
Drawer close
100
100
100
96
100
100
Drawer open
99 ∗
98 ∗
100 ∗
100
97
100
Appendix
TABLE A10: Per-session success (%) of the coding agents under the strict interface . Each session’s frozen program is evaluated on 100 hidden seeds. ∗ : the scored program solves inverse kinematics.
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.
Jiafeng Xu, Qi Li, Yan Shen +7
School of Computer Science, Peking University. · 1PrimeBot Research Institute, Swancor Advanced Materials Co., Ltd. · 3Crobotia.
Manipulating suspended payloads with humanoid robots is challenging because the robot can only influence an underactuated, oscillatory load through whole-body motion and intermittent contact. Imitation learning provides safe initial behavior but does not directly optimize final placement, while reinforcement learning from scratch is unsafe and sample-inefficient on real humanoids. We present HOIST-Humanoid Optimized with Imitation and Sample-efficient Tuning for manipulating suspended loads. HOIST first finetunes a high-level vision-language-action (VLA) policy from virtual-reality (VR) teleoperation demonstrations and executes its commands through a whole-body controller. It then uses VLA rollouts and iterative batched RL to improve placement accuracy and stopping behavior. Experiments in simulation and on a real humanoid show that HOIST improves over imitation-only and additional-demonstration baselines; compared with pure VLA rollouts, HOIST reduces translational placement error by 19.9 cm and raw angular error by 3.56 degrees, demonstrating the potential of humanoids for underactuated material-handling tasks.
Songyang Liu, Shunyu Yao, Dingyuan Huang +1
Department of Civil and Coastal Engineering University of Florida United States
Large Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulation remains challenging as the high-dimensional joint action space and tight inter-arm coordination constraints rapidly overwhelm standard context windows. To address this, we introduce BiCICLe (Bimanual Coordinated In-Context Learning), the first framework that enables standard LLMs to perform few-shot bimanual manipulation without fine-tuning. BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions. Evaluated on 13 tasks from the TWIN benchmark, BiCICLe achieves 70.5% average success rate, outperforming the best training-free baseline by 6.1 percentage points and surpassing most supervised methods. We also demonstrate superior real-world performance on 3 tasks without hardware-specific retraining. The project page is available at https://alesspalma.github.io/bicicle
Alessio Palma, Indro Spinelli, Vignesh Prasad +4
Sapienza University of Rome, Italy · TU Darmstadt, Germany · Hessian.AI, Germany