An LLM-powered Agent Framework for Heterogeneous Evacuation Behavior Modeling under a Moving Threat in a Public Plaza
Authors: Jian Ma, Runxin Yu, Tianyu Tang, Xiaolian Li
Organizations: School of Transportation and Logistics, Southwest Jiaotong University, Chengdu, 610031, China · Fujian Police College, Fuzhou, 350007, China
Modeling heterogeneous evacuation behavior under a moving threat is difficult because human perception, memory, and evidence evaluation are not well captured by fixed rules. We propose a novel LLM-powered agent-based framework to represent these internal decision processes. Each pedestrian agent perceives a private symbolic ASCII view, maintains a Memory-based Knowledge Graph derived solely from individual observations, and makes decisions through persona-conditioned prompts under a common sampling configuration. A compressed decision context with stateless memory preserves trial-and-error experience across turns while excluding reasoning traces, and a validation engine separates behavioral choice from physical feasibility by executing routes only over observed terrain. We evaluated eight personality compositions in eight paired randomized blocks within a simulated public plaza. Usable-exit knowledge was strongly associated with evacuation success: 89.5% of agents possessing such knowledge evacuated, compared with 1.05% of those without it. Personality compositions also differed in their evaluation of remembered threat evidence: the proportion of danger assessments varied by 0.265, while high-urgency, low-directness decisions ranged from 11.41% to 34.33%. After direct threat sightings, responses converged, with 99.6% of assessments classifying the situation as dangerous. Movement was selected in 99.8% of decisions. Overall, evacuation outcomes were strongly associated with information access, while evacuation time was jointly associated with spatial geometry, information, and affect. The framework provides an auditable approach to generating endogenous behavioral heterogeneity through persona-conditioned LLM agents in crowd-evacuation simulations.
Figures & tables
Figure 1: Perception, decision, memory, and physical execution in the LLM-powered framework. (a) Local observations update each pedestrian’s private known map Bit . (b) The decision model combines the current view Pit , personality, affect, short-term context Cit , and long-term memory Mit . (c) The model selects an action and target git ; the engine constructs and checks routes through observed, passable cells. Memory revisions are reconstructed as knowledge graphs after the run.
Archetype
O
C
E
A
N
Average
50
50
50
50
50
Calm-independent
50
55
40
45
30
Rule-following
40
70
45
60
50
Anxious-conformist
40
45
62
60
72
Bold-solitary
68
35
35
40
32
Table 1: FFM T-scores of the five personality archetypes.
Block
Content
Identity
An ordinary pedestrian acting on local perception, its own known map, memory, affect, and messages received
Map
The fixed north-up orientation, the @ -relative offset convention, and the glyph table of Table 3
Judgment
Target choice belongs to the agent, route legality to the engine; the evidence clause; the default target-distance band
Intent
intent and urgency as a standing goal that persists until the agent changes it
Route options
Fields of the engine-computed routes (length, familiarity, crowding, clearance from A or D , reach) and the rule that the shortest route need not be chosen
Past experience
Use of the decision-context capsule and of the engine’s feedback on the previous response
Table 2: The seven instruction blocks of the shared policy section.
Figure 2: Personality-to-prompt encoding in the controlled decision interface. (a) An anxious-conformist profile illustrates the five FFM T-scores and their position relative to T=50 . (b) The agent-specific identity section converts demographics and scores into bands, salient facets, and behavioral anchors. (c) The shared policy, tool schema, and temperature =0.2 remain fixed across agents, and the resulting identity section is inserted into every decision prompt.
Figure 3: From geometry to action. The filled view cone and first-order Moore neighborhood form the perception map, unknown cells appear as ? , and the private Memory-based Knowledge Graph supports engine-generated route options as well as A* validation of a model-selected target.
Symbol
Meaning
?
Outside the perception map or occluded
.
Walkable ground
#
Wall or outside the building
T
Tree or fixed obstacle
E
Exit
G
Gate, open
Table 3: Character vocabulary of the local symbolic view.
Figure 4: Recorded evolution of Agent 5’s long-term first-person memory in the completed GLM-5.3-flash reference run. Facts are arranged by version ( v1 – v8 ) and semantic track. Solid dark lineages were retained, blue nodes mark revisions, and pale dashed lineages terminate facts later omitted from memory. Dashed cross-track edges show recorded semantic associations, such as opening knowledge informing route intention and a heard shout accompanying an affect update. Reworded facts are matched by character-bigram Jaccard overlap at J≥0.40 . The final record contains six active and 18 forgotten facts; graph construction occurs after the simulation.
Figure 5: Plaza geometry and three illustrative trajectories. The perception map shows the three evacuation exits, two entrance gates, eleven pedestrian starting cells, fixed obstacles, and the initial threat position. The concentric circles indicate the implemented 2 m attack radius and 12 m perception radius. The trajectories illustrate three recorded outcomes from one completed run.
Measure
Definition
Evacuation
Evacuation time
Turns from run start to reaching an opening
Trajectory length
Distance walked along the executed trajectory
Evacuation efficiency
Shortest exit distance over walked distance, capped at 1; zero when unresolved
Critical safety margin
Smallest distance to the adversary over the run
Mean threat distance
Mean distance to the adversary over the run
Table 4: Outcome measures computed for each pedestrian.
Figure 6: Decision revision by layer and stimulus. Left: the frequencies of goal-, plan-, and execution-level revisions across 10,947 decisions, together with the proportion retaining the previous decision. Right: cumulative revision probabilities for six stimulus classes. The dashed line marks the 4.3% goal-revision rate for epochs with an absolute urgency change below 0.15.
Stage
Memory excerpt (translated)
Interpretation
Decision or plan
Case 1: From uncertain signals to escape and exit search
1. Situation unclear
“I wanted to stop and observe to clarify the situation, without hastily following the others.”
The meaning of others’ behavior is still unclear.
The stated plan is to observe before deciding whether to follow others.
2. Threat confirmed
“The 2 shouts I heard earlier … were probably warnings.” “Escaping comes first; exploration can wait.”
Earlier sounds are now interpreted as warnings.
The decision changes from continuing activity to moving away from the threat.
3. Another exit needed
“Next, I need to find an exit that is still open or a direction with more people, and keep moving away from the northeast.”
Escape remains necessary, but the known gate may be unusable.
The next decision maintains threat avoidance and includes looking for other exits.
Case 2: From choosing a known exit to checking and abandoning it
1. Known exit chosen
“I remembered that [the gate] was to the south, so I decided to go south, away from the armed person and toward the known exit.”
The known exit provides a destination for escape.
The selected goal is to reach the known exit.
Table 5: Changes in remembered information and escape plans in 2 illustrative cases.
Figure 7: Reported urgency and subsequent movement directness across 10,947 decisions. (a) Dashed lines mark the pooled medians of urgency ( u=0.60 ) and directness ( d^=0.606 ); quadrant labels give the percentages of all evaluated decisions. Symbols mark outcome-group means, with each pedestrian weighted equally. (b) The 5 homogeneous personality compositions. Values above each panel give the mean run-level share of high-urgency, low-directness decisions and its 95% confidence interval from paired-block resampling. In all panels, the linear color scale shows the percentage of decisions within that panel falling in each cell. The individual share of high-urgency, low-directness decisions distinguished death from evacuation with an AUC of 0.839.
Figure 8: The recorded chain for Agent 3 in the rule-following run. Five lanes carry perception events, the affect transition and the self-reported urgency trace, the appraisal band, each decision at a height set by its revision depth, and the write, rewrite, and forget operations on memory. Callouts reproduce selected memory entries. The pedestrian learns of only one opening, the south entrance gate, which closes at 11.67 s.
Figure 9: Exit knowledge and evacuation outcomes. (a) Outcome shares by the number of usable exits known. (b) Outcome shares for the 4 exit-information groups. “No initial exit information” refers to the absence of system-provided exit information at initialization; “Only a closed gate reported” refers to information received from others during the simulation. Numbers beside the bars give group sizes ( n ) and evacuation rates. (c) Cumulative distributions of time to first usable-exit knowledge among pedestrians who acquired it, grouped by final outcome. Markers indicate median times of 5.17 s for evacuated pedestrians and 32.33 s for killed pedestrians.
Figure S1: Shout and exit-information propagation in one run. The same eleven pedestrians hold identical positions in both panels and are encoded by outcome. Directed edges give deduplicated sender–receiver pairs for (a) heard shouts and (b) transmitted exit information. Across the experiment, shout out-degree carries an AUC of 0.808 for being killed and exit knowledge an AUC of 0.651 for rescue.
Figure S2: Reconstructed knowledge graph for Agent 3. Nodes are grouped into eleven type clusters and relations are bundled between types, with RETRIEVED FOR emphasized. This agent’s graph holds 54 observations, 54 decisions, 54 routes, 163 memories, and 12 affect records, with an average of 6.67 memory links per decision. Across the experiment, the graphs contain 75,944 nodes and 138,963 edges.