Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
Figures & tables
Figure 1: (a) An illustration of DuoMind , a distributed framework combining an orchestrator and action model for multi-robot coordination. (b) RoboPoly is a benchmark we develop for long-horizon coordination under distributed robot control. (c) DuoMind consistently outperforms the baselines across both benchmarks.
Figure 2: In DuoMind , each robot operates with a VLM as its orchestrator for high-level reasoning and a VLA as its action model for low-level control. At each high-level planning step, the orchestrator at each robot reasons and generates both a low-level instruction for the action model and a structured message for other robots. The action model subsequently follows the low-level instruction and produces fine-grained actions.
Method
Hang Bag
Food Serve
Prepare Snack
Clean Table
DuoMind w/ π0.5
78.00
32.75
18.00
24.75
π0.5 only
52.25
29.75
16.25
16.50
π0 only
58.75
22.25
3.75
0.00
Table 1: Comparison of success rates on RoboPoly tasks.
Method
Hanging Mug
Pick Diverse Bottles
Stack Bowls Two
Put Bottle Dustbin
DuoMind w/ π0.5
24.75
58.50
85.50
34.50
π0.5 only
17.25
48.50
82.50
31.25
π0 only
12.50
42.50
33.75
21.25
Table 2: Comparison of success rates on RoboTwin tasks.
Figure 3: We observe that explicit inter-agent communication enables the robots to better coordinate their behaviors, such as exchanging bread sequentially or placing objects on different plates. By avoiding conflicting or redundant actions in a narrow shared workspace, the robots achieve more efficient and safer coordination.
Figure 4: Ablation studies of DuoMind . (a) Success rates when inter-agent communication is disabled. (b) Success rates with different action models.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The task overview of RoboPoly .
Figure 6: The task overview of selected RoboTwin tasks.
Figure 7: The result of ablation study that disable the inter-agent communication in DuoMind .
Figure 8: Results of the ablation study using different action models in DuoMind .
Task
Method
Train steps
Batch
Peak LR
Execute / horizon
VLM interval
Hang Bag
π0.5 / π0 only
10k
256
2.5×10−5
12/16
–
Hang Bag
DuoMind
10k
256
2.5×10−5
12/16
4
Food Serve
π0.5 / π0 only
10k
256
2.5×10−5
16/16
–
Food Serve
DuoMind
10k
256
2.5×10−5
16/16
4
Prepare Snack
π0.5 / π0 only
2k
512
1×10−4
16/16
–
Prepare Snack
DuoMind
2k
512
1×10−4
16/16
4
Appendix
Table 3: Fine-tuning and evaluation settings on RoboPoly .
Task
Method
Train steps
Batch
Peak LR
Execute / horizon
VLM interval
Hanging Mug
π0.5 / π0 only
20k
32
2.5×10−5
50/50
–
Hanging Mug
DuoMind
20k
32
2.5×10−5
50/50
4
Pick Diverse Bottles
π0.5 / π0 only
20k
32
2.5×10−5
50/50
–
Pick Diverse Bottles
DuoMind
20k
32
2.5×10−5
50/50
4
Stack Bowls Two
π0.5 / π0 only
20k
32
2.5×10−5
50/50
–
Stack Bowls Two
DuoMind
20k
32
2.5×10−5
50/50
2
Appendix
Table 4: Fine-tuning and evaluation settings on RoboTwin.
Figure 9: We evaluate DuoMind on 7 RoboPoly tasks.
Figure 10: We evaluate DuoMind on 8 selected RoboTwin tasks.