STRATA: Self-Learning Through Role-Aligned Tiered Agents for Real-Time Strategy Games
Organizations: School of Instrumentation Science and Optoelectronic Engineering, Beijing University of Aeronautics and Astronautics · Tsinghua University · University of Chinese Academy of Sciences · Department of Automation, Tsinghua University · CFINS, Tsinghua University
Abstract
Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models to command decision-making in RTS games, enabling agents to read textual game states and generate high-level plans. However, long inference latency can cause them to miss critical tactical events. The complexity and tactical diversity of full RTS matches also leave existing systems heavily dependent on manually written experience-based prompts, with limited ability to learn continuously from past games. We present STRATA, a role-aligned hierarchical system with cross-game self-learning for Red Alert. STRATA assigns in-game strategic, logistical, and tactical decisions to a Strategic Agent (SA), Logistics Agent (LA), and Tactical Agent (TA), respectively. The SA generates high-level directives based on the global game state and relevant experience cards, while the LA and TA handle logistics and tactical execution. After each match, a Review Agent (RA) derives candidate experience from game traces, validates and revises it using evidence from subsequent matches, and compresses strategic experience supported across multiple games into concise experience cards for SA retrieval. We evaluate STRATA through the formation of experience cards, full-match comparisons before and after learning, and experience learning against AI opponents with different play styles. Under a fixed scenario, using the learned experience cards increases the observed win rate from 30% to 100%. Sequential learning against AI opponents with different play styles also produces distinct long-term strategic experience.
Figures & tables
| Experience theme | Rush condition | Turtle condition |
|---|---|---|
| Vehicle production | Lower the production threshold under a thin economy. With fewer than four harvesters, allow cash to fall to about 100 and queue one vehicle at a time to prevent the war factory from idling. | Use stage-dependent cash floors. Keep about 200 before the second refinery. Afterward, produce vehicles while cash exceeds 1,500 and stop near 800. |
| Economic and production expansion | Prioritize the second refinery over additional production capacity, radar, and technology to stabilize the basic economy before further development. | Under low pressure, prioritize a third refinery and another war factory so that economy and production capacity expand together. |
| Post-technology attack | Produce advanced units in short batches and wait for an attack window. | Reduce prolonged rallying and waiting. When at least six combat units are available and few opponents are visible, the SA is more likely to issue a proactive attack directive. |