Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a benchmark for \emph{Arm-wise Compositional Generalization} that provides a common testbed for studying skill recomposition in dual-arm policies. It contains 23 task--condition pairs across 8 task families, with 6 in-domain conditions and 17 unseen compositions covering reordering, synchronization, their combination, and cross-task composition. All methods receive the same per-arm atomic prompts, and success requires achieving the task goal while satisfying physical milestones and specified order or timing constraints. Using π0.5 as a common vision-language-action backbone, we compare representative data-augmentation and architectural strategies with shared source data and a common evaluation protocol. Our architectural study examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure. Combining these choices yields \textbf{AE-VLA}, which achieves 21.53% generalization success in simulation, compared with 2.94% for Single π0.5, 3.06% for MA-VLA, and 5.53% for two independently controlled π0.5 policies. On physical SO101 robots, AE-VLA reaches 39.00% mean success across five unseen conditions, compared with 10.00% for the strongest baseline. These findings provide empirical guidance for designing dual-arm policies that generalize beyond fixed training routines.
Figures & tables
Figure 2: Three design choices within one π0.5 policy. (a) Single uses H joint action tokens to predict both arms. Token Group uses two groups of H tokens, each assigned to one arm, within a shared action expert. (b) SkillLoRA selects a skill adapter for each group. (c) AWA controls attention between groups while retaining global context. Here H=50 ; arrows in (a) summarize decoding and denoising.
Figure 3: Success and progress by test group. Parentheses give condition counts; each condition has 100 episodes per method. TG means Token Group. Bars show means over conditions. Generalization averages all 17 non-in-domain conditions.
Task
Condition
Single π0.5
MA-VLA
Dual π0.5
AE-VLA
Stack Bowls
In-domain
74.00
77.00
83.00
83.00
Reorder
0.00
7.00
4.00
79.00
Sync
0.00
0.00
0.00
26.00
Sync+Reorder
0.00
0.00
1.00
67.00
Stack Cubes
In-domain
17.00
28.00
27.00
20.00
Reorder
0.00
1.00
9.00
19.00
Table 1: All simulation conditions: CCSR (%). Each condition has 100 episodes per method. Only Dual π0.5 uses two policies. Parentheses in cross-task rows retain the timing setting. Bold marks the best result.
Variant
Token Group
SkillLoRA
Arm-wise Attn
ID CCSR
Gen. CCSR
Gen. Progress
Single π0.5
–
–
–
37.00
2.94
10.94
Token Group
✓
–
–
45.00
3.82
10.98
Token Group + SkillLoRA
✓
✓
–
41.33
6.00
19.02
Token Group + AWA
✓
–
✓
25.00
5.82
28.84
Token Group + Both (AE-VLA)
✓
✓
✓
27.17
21.53
44.25
Table 2: Design choices on the same π0.5 backbone. The four grouped variants form a 2×2 comparison of SkillLoRA and AWA. Values are percentages averaged over conditions.
Figure 4: SO101 examples. Rows show Stack Bowls/In-domain, Stack Bowls/Reorder, Stack Bowls/Sync, and Cube in Bowl/Cross-task. Table 3 reports all evaluation trials.
Setting
Task
Single π0.5
MA-VLA
Dual π0.5
AE-VLA
In-domain
Stack Bowls
15/20
15/20
18/20
17/20
Stack Cubes
5/20
3/20
10/20
9/20
Push Cubes
10/20
9/20
14/20
10/20
ID mean (%)
50.00
45.00
70.00
60.00
Reorder
Stack Bowls
0/20
4/20
0/20
9/20
Reorder
Push Cubes
0/20
2/20
0/20
6/20
Table 3: SO101 results with distractor objects. Each task row shows successes out of 20 trials. Mean rows average rates over three in-domain or five unseen conditions. Bold marks the best result.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Task family
ID
Reorder
Sync
Sync+R
Cross-task
Stack Bowls
✓
✓
✓
✓
–
Stack Cubes
✓
✓
✓
✓
–
Push Cubes
✓
✓
✓
✓
–
Burger & Fries
✓
✓
✓
✓
–
Can in Basket
✓
–
✓
–
–
Cube in Bowl
–
–
–
–
Reference, Sync
Appendix
Table 4: ACG-Bench condition inventory. Reference and Sync distinguish the execution requirements within the Cross-task group.
Skill
Behavior and arguments
Example atomic prompt
Pick
Approach and grasp the specified object.
Pick up the bowl.
Place
Carry a held object to the target and release it.
Place the can into the basket.
Lift
Raise a grasped object or container to the required pose.
Lift the basket.
Pull
Move a grasped articulated part to open it.
Pull the drawer open.
Move
Position the end effector for the next manipulation.
Move to the blue cube.
Push
Move an object toward a target through contact.
Push the blue cube to the target.
Appendix
Table 5: Atomic operations and example prompts. Object and target arguments vary across tasks. These semantic names do not specify numerical adapter IDs.
Task family
Manipulation operations
Source behavior or held-out combination
Stack Bowls
Pick, Place
Each arm handles a bowl and places it at its target.
Stack Cubes
Pick, Place
Each arm handles a cube and places it at its target.
Push Cubes
Move, Push
Each arm approaches a cube and pushes it to a target.
Burger & Fries
Pick, Place
The arms place the burger and fries at their respective targets.
Can in Basket
Pick, Lift, Place
One arm holds and lifts the basket; the other places the can inside.
Object in Cabinet
Handle grasp, Pull, Pick, Place
One arm grasps and opens the drawer; the other places the box inside.
Appendix
Table 6: Manipulation operations and their roles in ACG-Bench. Source rows summarize demonstration prompts; the final two rows describe held-out compositions. Wait and recovery behavior are discussed in Appendix A.3 .
Phase
Left-arm prompt
Right-arm prompt
1
Pick up the blue cube.
Wait for the other robots.
2
Put the blue cube on the target.
Wait for the other robots.
3
Recover.
Pick up the green cube.
4
Wait for the other robots.
Place the green cube on the target.
5
Wait for the other robots.
Recover.
Appendix
Table 7: The five prompt pairs in the Stack Cubes source conversion. The two prompt strings are stored together and separated by arm at policy input. This is a training annotation, not a fixed test-time clock.
Test group
Change to the source routine
Reorder
Require kR≺dR≺kL≺dL : the right arm completes its manipulation first.
Sync
Require ∣τ(kL)−τ(kR)∣≤δ , followed by the source placement order dL≺dR .
Sync+Reorder
Use the same synchronized pick pair, then reverse the placement order to dR≺dL .
Cross-task
Combine bowl handling from Stack Bowls with cube placement from Stack Cubes; the cube’s target is now a bowl.
Appendix
Table 8: How a familiar routine can be recomposed. The first three rows illustrate changes in order and timing; the last illustrates source-skill reuse across tasks. Exact milestones remain task-specific.
Component
Setting
Initialization
π0.5 base checkpoint
Backbone / action expert variants
gemma_2b / gemma_300m
Control dimensions / model dimensions
2×8 / 32 (zero-padded)
Action horizon / grouped action tokens
50 / 2×50
Prompt/state token budget
200 per arm; 400 in total
Optimizer
AdamW; (β1,β2)=(0.9,0.95)
Appendix
Table 9: Multitask AE-VLA training configuration.
Task
Condition
TG
TG + SL
TG + AWA
AE-VLA
Stack Bowls
In-domain
88.00
83.00
43.00
83.00
Reorder
0.00
32.00
0.00
79.00
Sync
0.00
19.00
0.00
26.00
Sync+Reorder
0.00
17.00
7.00
67.00
Stack Cubes
In-domain
34.00
34.00
8.00
20.00
Reorder
0.00
3.00
1.00
19.00
Appendix
Table 10: Four design settings on all conditions: CCSR (%). TG means Token Group; SL means SkillLoRA. Each setting uses one policy and the same 23 conditions, with 100 episodes per condition.
Figure 5: SO101 setup. Two arms share a workspace observed by one fixed global camera and one wrist camera on each arm.
Figure 6: SO101 training demonstrations. Example observations from the global, left-wrist, and right-wrist cameras (left to right). We collect 50 demonstrations per training task through LeRobot teleoperation.