An agent that remembers what it did on a web page must decide when two pages count as the same. Memories built on observation similarity merge pages that look alike but behave differently, and GUIs are full of such pages: two tabs of one widget or two rows of one menu answer the same click differently. We define the merge rule as an action-conditioned bisimulation over the empirical predictive state graph a frozen agent fills as it acts. Two states merge only when their shared actions lead to agreeing outcomes and successor blocks under an affordance label. Observation similarity never enters the rule, and nothing is trained. It replaces the merge rule of an existing outcome-value memory, so a closed-loop comparison isolates it. On MiniWoB++ it raises success rate over a memoryless agent, while a control taking identical exploratory detours, the prior successor-representation merge, and the same criterion without action conditioning change nothing.
Figures & tables
Figure 1: The loop the criterion sits in. A frozen policy acts on the GUI and every transition is recorded, so each state accumulates an empirical profile of what its tried actions led to. States are compared only through those profiles, never through what the pages look like, and iterating the comparison to stability gives the partition B . Everything to the right of B is inherited unchanged from the outcome-value memory the criterion is dropped into.
Figure 2: One evaluation cell of click-tab-2 (goal “Switch between the tabs to find and click on the link gravida. ”). The panels are the three states reachable before the target is visible; only (c) holds the link. They differ by a highlight and a paragraph of filler, so an observation-similarity criterion merges them, while their responses to the same three clicks differ and ( 4 ) separates them. Below, the element clicked at each step. ReAct alternates Tab #2 and Tab #1 for all eight steps. APSG emits the same five actions, then overrides at step 6 (bold), once the block holds enough recorded failures of click(18) for ( 8 ) to prefer click(22) .
System
SR
Δ vs. ReAct
95% CI
p
Memoryless
ReAct
0.723
–
–
–
Shortlist control
0.723
+0.000
[+0.000,+0.000]
1.000
Memory, prior merge rules
SR merge
0.728
+0.005
[−0.007,+0.018]
0.727
Observation merge
0.708
−0.015
[−0.033,+0.003]
0.180
Table 1: Closed-loop MiniWoB++ over 10 tasks with a frozen Qwen3-8B, single seed. SR is the benchmark’s own success rate on the evaluation half fixed before the run (400 paired cells per system), decoded greedily by every system. Environment seeds are identical across systems, so Δ against ReAct is paired cell by cell, with a 95% percentile bootstrap interval and an exact McNemar p .
Figure 3: Both panels recomputed from the episode log. (a) Cumulative success rate, with a 95% bootstrap band on our curve only. The shortlist control takes the identical detours but keeps nothing, so it recovers to ReAct’s rate and stops there; under the action-conditioned partition the same detours are paid off by episode 35 . (b) On the evaluation episodes: evidence coverage , the share of decisions whose block held recorded evidence for a candidate; override rate , the share at which ( 8 ) displaced the policy’s first choice; override precision , the share of displaced outcomes that moved from failure to success rather than the reverse.