This work examines the transfer of a co-evolved communication mechanism between two robotic agents from a discrete two-dimensional (2D) simulator to a three-dimensional simulator with real physics (3D). The study focuses on whether a communication mechanism co-evolved in a 2D environment retains its functional role after transfer to a 3D physics-based simulator. To support this analysis, the effects of the episode time budget, the social cue, and the asymmetry between the two co-evolved roles were examined. The results indicate that the success rate increased approximately linearly with the evaluated time budgets, with no evidence of a plateau between 2,000 and 6,000 physics steps, suggesting that evaluations based on shorter episodes may underestimate the performance of the trained controllers. In both simulators, the social cue functioned primarily as a jam- assistance mechanism rather than as a navigation guide, although with a more pronounced effect in 2D. Analysis of eight independent evolutionary runs revealed a consistent direction of asymmetry, although its magnitude varied across runs. Controlling the processing order between agents allowed us to rule out an artifact of the physics engine. Finally, the results are discussed in terms of the factors that may contribute to the remaining performance gap observed after transfer.
Figures & tables
Figure 1: Comparison of the simulation engines used for the evaluation of coevolved communication between robotic agents. (a) 2D engine with discrete environment representation, without physical modeling, used for low computational-cost simulations. (b) 3D engine based on PyBullet, with 240 Hz physical simulation and decoupled simulation time control through decision updates every 15 physics steps ( ≈ 16 Hz), equivalent to a control cycle commonly used in robotic simulators. Both environments contain two agents, shared obstacles, and shared objectives for performing cooperative tasks.
Figure 2: Distribution of the robot’s eight proximity sensors. The sensors are evenly spaced at 45 ∘ intervals around the robot, providing 360 ∘ detection coverage. The black line indicates the robot’s forward-facing direction and serves as a visual reference for interpreting sensor orientation in both the 2D and 3D simulators.
Figure 3: Success rate of agents A and B as a function of the episode time budget (2,000 to 6,000 physics steps) for checkpoint v15_full_final without retraining. Error bars represent 95% confidence intervals for a binomial proportion ( n=300 episodes per budget). The success rate increased for both agents as the time budget increased. Agent B maintained a higher success rate than agent A across all evaluated budgets.
Budget
With signal
Without signal
Difference
n
p
(steps)
(%)
(%)
(pp)
(with signal)
2,000
11.1
3.5
7.6
36
0.146
3,000
11.2
2.3
8.8
269
<0.0001
4,000
9.6
2.0
7.6
1,020
<0.0001
6,000
12.2
1.5
10.7
2,054
<0.0001
Table 1: Stall resolution rate with and without social signal, by episode time budget.
Figure 4: Stall resolution rate with social signal present (heard) and absent (silent), as a function of the episode time budget. Error bars show the 95% confidence interval for a binomial proportion, calculated from the counts observed in each condition. For the 2,000-step budget, the observed difference did not reach statistical significance, coinciding with the small number of decision points classified as stalled with a signal present ( n=36 ).
Figure 5: B/A weight ratios in the right_turn action column ( w_out and w_res ) for eight independent evolutionary runs (four in the 3D simulator and four in the 2D simulator). The horizontal dashed line indicates parity between agents (ratio = 1.0). The w_res ratio was significantly above parity ( p=0.012 ), whereas the w_out ratio showed the same trend without reaching conventional statistical significance ( p=0.098 ). The labels on the x-axis correspond to the original checkpoint names of the evolutionary runs and are retained for traceability.
This work evaluates the direct transfer of a co-evolved communication protocol from a 2D simulation to a 3D physical environment, without retraining the network weights. Two e-puck-type robots, controlled by a GRU network with residual connection, were evaluated in a food-seeking task with social signaling. The sensory and motor translation layer required three corrections for stable physical operation, including the calibration of a hunger term based on a measurable asymmetry in the trained residual weights. Even with these corrections, the transfer was partial and asymmetric: one agent reached the food source in one of thirty tested seeds, while the other did not reach it in any. Task success was measured by both agents reaching the food area. An additional experiment incorporating explicit directional information in the social channel produced observable changes in the trajectory of the receiving agent and improvements in several specific cases. However, these improvements were not enough to allow the second agent to reach the food source, suggesting that the limitation may not be explained solely by signal translation, but also by the ability to navigate under the new physical constraints. The results suggest that successful transfer of emergent communication may depend not only on preserving the signaling process itself, but also on preserving the ecological and navigational conditions under which the protocol evolved.
Self-improving language agents are typically evaluated in isolation: an agent attempts a task, receives feedback, and iteratively refines its own behavior. Yet agents increasingly operate alongside peers whose strategies and outcomes are publicly visible. This raises an under-studied question: when does shared experience produce improvements that self-improvement alone cannot achieve? We introduce SAGE (Social Agent Group Evolution),an evaluation framework that compares two compute-matched conditions: SocialEvo, where agents from five distinct model families co-evolve with access to all peers' histories; and SelfEvo, where each agent receives the same number of task attempts but sees only its own past, which is conventional in self-improving agent studies. We instantiate SAGE in three arenas: open-ended ML research, long-horizon economic planning, and strategic multiplayer play, evaluated across multiple evolutionary rounds. We find that group history is not a universal amplifier: the strongest agent does not exceed its self-evolution ceiling. However, agents that plateau under self-improvement can achieve significant breakthroughs when peer experience is available. In competitive settings, counterfactual controls reveal that agents improve generally rather than developing opponent-specific strategies. Across different forms of shared history, filtered peer traces and reflective summaries often outperform raw logs, indicating that social gains depend on abstraction rather than exposure volume. These findings reveal that peer-history gains are agent-specific, arena-dependent, and contingent on the capacity to abstract transferable knowledge from public traces.
In recent years, reinforcement learning (RL) has shown remarkable success in robotics when a fast and accurate simulator is available for a given task. When using RL and simulation, more simulator realism is generally beneficial but becomes harder to obtain as robots are deployed in increasingly complex and widescale domains. In such settings, simulators will likely fail to model all relevant details of a given target task and this observation motivates the study of sim2real with simulators that leave out key task details. In this paper, we formalize and study the abstract sim2real problem: given an abstract simulator that models a target task at a coarse level of abstraction, how can we train a policy with RL in the abstract simulator and successfully transfer it to the real-world? Our first contribution is to formalize this problem using the language of state abstraction from the RL literature. This framing shows that an abstract simulator can be grounded to match the target task if the grounded abstract dynamics take the history of states into account. Based on the formalism, we then introduce a method that uses real-world task data to correct the dynamics of the abstract simulator. We then show that this method enables successful policy transfer both in sim2sim and sim2real evaluation.
Yunfu Deng, Yuhao Li, Josiah P. Hanna
Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI 53706 USA · Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Amherst, MA 01003 USA