Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters. The logic path receives only the core turn and emits a verifiable structured state (Micro-State); the persona path renders it in character with the full history. In same-base-model ablations on an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs over 4 pollution levels x 3 arms x 2 tasks x 2 personas x 5 seeds) we find: (i) the decoupled logic path is structurally invariant to pollution: its prompt stays at 180 (Llama) or 167 (Gemma) tokens while the mixed single-pass prompt grows from 242 to 1,203, and its outputs are byte-identical across levels (40/40); (ii) the mixed single pass degrades monotonically (composite logic score 0.669 to 0.150 on Llama, 0.487 to 0.150 on Gemma), mostly by failing to emit the required structured output (80-95% of runs on Llama, 100% on Gemma at the two highest levels); (iii) with the same pollution fed into the decoupled logic path, the dedicated-adapter, dedicated-format path is still more robust than the single pass on the 8B model (failure 0-20% vs 80-95%; paired Δ +0.30 to +0.50, Cliff's δ 0.50-0.85, Holm-adjusted p≤0.03) but not on the 4B model, where both collapse. Separation costs one extra decode on a topic's first turn (28.2 s vs 18.2 s on Llama) and buys persona hot-swapping in 1.7 ms without re-running the logic path. Code, rubric, fixtures, adapters and logs are released.
Figures & tables
Figure 1: Figure 1. AO-DA data flow. The logic (What) path receives a bounded, history-free input and emits a schema-checked Micro-State; the persona (How) path renders the Micro-State in character while receiving the full history. Both paths run as adapter swaps on one resident INT4 base. Dashed lines mark the Micro-State cache and the adapter hot-swap.
Arm
What path input
How path input
Stages
Role
G0 mixed
— (single pass: logic header + persona overlay + full pollution)
—
1
Baseline: What and How in one context
G3 decoupled
core turn only (clean)
Micro-State + full pollution
2
The architecture
G3-P decoupled, logic polluted
core turn + full pollution
Micro-State + full pollution
2
Ablation: separation without the input bound
Table 2
Level
Chat history
Long user input
System bloat
Note
L0
—
—
—
Baseline, identical to the low-pollution Run 1 condition
L2
8–10 turns, domain-specific
—
—
Misleading, persona-heavy
L3
as L2
≈2,000 chars appended
—
L4
as L2
as L3
persona few-shot examples appended to the system prompt (621–811 chars)
Table 3
Arm
n
Composite logic (mean ± SD)
How coverage
Failure rate
Latency s
G0 mixed single pass
30
0.588 ± 0.253
0.378
6.7%
16.0
G3 decoupled
30
0.583 ± 0.131
0.633
0.0%
26.2
G4 decoupled + constraint echo
30
0.583 ± 0.131
0.711
0.0%
30.2
Persona LoRA on logic path (probe)
30
0.524 ± 0.240
—
13.3%
11.5
Table 4
Base
Arm
L0
L2
L3
L4
Llama-8B
G0
0.669 ± 0.230
0.450 ± 0.266
0.219 ± 0.172
0.150 ± 0.000
Llama-8B
G3
0.655 ± 0.080
0.655 ± 0.080
0.655 ± 0.080
0.655 ± 0.080
Llama-8B
G3-P
0.655 ± 0.080
0.798 ± 0.137
0.515 ± 0.319
0.650 ± 0.258
Gemma-4B
G0
0.487 ± 0.199
0.209 ± 0.162
0.150 ± 0.000
0.150 ± 0.000
Gemma-4B
G3
0.695 ± 0.234
0.695 ± 0.234
0.695 ± 0.234
0.695 ± 0.234
Gemma-4B
G3-P
0.695 ± 0.234
0.182 ± 0.100
0.150 ± 0.000
0.150 ± 0.000
Table 5
Figure 2: Figure 2. Composite logic score by pollution level (mean, 95% bootstrap CI, n = 20). The decoupled arm (G3) is identical across levels by construction. Horizontal dotted line: scoring floor (empty structured output).
Base
Arm
Lv
Fail % [95% CI]
Raw-text
Speech
How cov.
Llama-8B
G0
L0
0 [0, 16]
0.779 ± 0.172
0.613 ± 0.207
0.500
Llama-8B
G0
L2
25 [11, 47]
0.713 ± 0.193
0.524 ± 0.243
0.467
Llama-8B
G0
L3
80 [58, 92]
0.328 ± 0.251
0.269 ± 0.194
0.083
Llama-8B
G0
L4
95 [76, 99]
0.380 ± 0.291
0.338 ± 0.250
0.167
Llama-8B
G3
L0
0 [0, 16]
0.655 ± 0.080
0.680 ± 0.311
0.633
Llama-8B
G3
L2
0 [0, 16]
0.655 ± 0.080
0.782 ± 0.213
0.783
Table 7
Figure 3: Figure 3. Top: structured-output failure rate (bars; whiskers are Wilson 95% CIs, so a 0/20 cell shows an upper bound of 16%). Bottom: raw-text logic score (format-independent). Under misleading persona-heavy history the mixed single pass mostly stops emitting the schema; the content loss is smaller but real.
Base
Comparison
Failures
Fisher p
Holm p
Llama-8B
G0 vs G3-P fail @ L0
0/20 vs 0/20
1.000
1.000
Llama-8B
G0 vs G3-P fail @ L2
5/20 vs 0/20
0.047
0.141
Llama-8B
G0 vs G3-P fail @ L3
16/20 vs 4/20
<0.001
0.001
Llama-8B
G0 vs G3-P fail @ L4
19/20 vs 3/20
<0.001
<0.001
Llama-8B
G0 fail L0 vs L2
0/20 vs 5/20
0.047
0.141
Llama-8B
G0 fail L0 vs L3
0/20 vs 16/20
<0.001
<0.001
Table 9
Base
Arm
Metric
Spearman ρ
p
n
Llama-8B
G0
composite
-0.737
<0.001
80
Llama-8B
G0
raw
-0.555
<0.001
80
Llama-8B
G0
failure
+0.760
<0.001
80
Llama-8B
G3-P
composite
+0.019
0.868
80
Llama-8B
G3-P
raw
+0.076
0.505
80
Llama-8B
G3-P
failure
+0.257
0.021
80
Table 10
Base
Lv
G3-P
G0
Δ [95% CI]
δ
p
Holm p
Llama-8B
L0
0.655
0.669
-0.014 [-0.11, 0.08]
+0.01
0.876
1.000
Llama-8B
L2
0.798
0.450
+0.347 [0.22, 0.46]
+0.72
<0.001
0.003
Llama-8B
L3
0.515
0.219
+0.296 [0.14, 0.45]
+0.50
0.006
0.028
Llama-8B
L4
0.650
0.150
+0.500 [0.39, 0.60]
+0.85
<0.001
0.001
Gemma-4B
L0
0.695
0.487
+0.207 [0.10, 0.32]
+0.55
0.002
0.013
Gemma-4B
L2
0.182
0.209
-0.026 [-0.09, 0.03]
-0.06
0.465
1.000
Table 11
Base
Lv
G3-P
G0
Δ [95% CI]
δ
p
Holm p
Llama-8B
L0
0.655
0.779
-0.124 [-0.20, -0.05]
-0.46
0.006
0.043
Llama-8B
L2
0.818
0.713
+0.105 [0.01, 0.20]
+0.33
0.044
0.220
Llama-8B
L3
0.678
0.328
+0.350 [0.21, 0.48]
+0.65
<0.001
0.007
Llama-8B
L4
0.662
0.380
+0.283 [0.08, 0.46]
+0.53
0.021
0.126
Gemma-4B
L0
0.800
0.731
+0.069 [0.01, 0.14]
+0.28
0.072
0.289
Gemma-4B
L2
0.563
0.691
-0.129 [-0.27, 0.02]
-0.35
0.083
0.289
Table 12
Base
Lv
G3
G0
Δ [95% CI]
δ
p
Holm p
Llama-8B
L0
0.655
0.669
-0.014 [-0.11, 0.08]
+0.01
0.876
0.876
Llama-8B
L2
0.655
0.450
+0.205 [0.08, 0.33]
+0.35
0.014
0.027
Llama-8B
L3
0.655
0.219
+0.436 [0.36, 0.50]
+0.86
<0.001
<0.001
Llama-8B
L4
0.655
0.150
+0.505 [0.47, 0.54]
+1.00
<0.001
<0.001
Gemma-4B
L0
0.695
0.487
+0.207 [0.10, 0.31]
+0.55
0.002
0.007
Gemma-4B
L2
0.695
0.209
+0.486 [0.37, 0.59]
+0.93
<0.001
<0.001
Table 13
Base
Arm
Lv
Logic prompt tokens [min–max]
Persona prompt tokens
Latency s (± SD)
Logic s
Persona s
Llama-8B
G0
L0
242 [232–252]
—
18.2 ± 4.4
18.2
0.0
Llama-8B
G0
L2
562 [498–625]
—
14.6 ± 4.1
14.6
0.0
Llama-8B
G0
L3
1,014 [950–1,077]
—
13.4 ± 7.2
13.4
0.0
Llama-8B
G0
L4
1,203 [1,115–1,291]
—
17.8 ± 11.4
17.8
0.0
Llama-8B
G3
L0
180 [179–181]
312
28.2 ± 10.3
11.8
16.4
Llama-8B
G3
L2
180 [179–181]
632
21.4 ± 4.3
12.1
9.4
Table 14
Base
Arm
Lv
Comp. Goku
Comp. Makima
How Goku
How Makima
Speech Goku
Speech Makima
Llama-8B
G0
L0
0.720
0.617
0.567
0.433
0.677
0.548
Llama-8B
G0
L2
0.340
0.560
0.367
0.567
0.445
0.603
Llama-8B
G0
L3
0.150
0.287
0.067
0.100
0.265
0.273
Llama-8B
G0
L4
0.150
0.150
0.167
0.167
0.338
0.338
Llama-8B
G3
L0
0.655
0.655
0.767
0.500
0.823
0.537
Llama-8B
G3
L2
0.655
0.655
0.900
0.667
0.853
0.713
Table 15
Base
Lv
G3
G0
Δ [95% CI]
δ
p
Holm p
Llama-8B
L0
0.633
0.500
+0.133 [-0.03, 0.32]
+0.23
0.016
0.082
Llama-8B
L2
0.783
0.467
+0.317 [0.10, 0.52]
+0.57
0.008
0.050
Llama-8B
L3
0.383
0.083
+0.300 [0.15, 0.47]
+0.42
0.007
0.050
Llama-8B
L4
0.600
0.167
+0.433 [0.23, 0.63]
+0.62
0.004
0.028
Gemma-4B
L0
0.400
0.450
-0.050 [-0.22, 0.12]
-0.11
0.832
1.000
Gemma-4B
L2
0.600
0.450
+0.150 [-0.00, 0.30]
+0.28
0.084
0.252
Table 16
Base
Lv
G3
G0
Δ [95% CI]
δ
p
Holm p
Llama-8B
L0
0.680
0.613
+0.067 [-0.07, 0.20]
+0.22
0.531
1.000
Llama-8B
L2
0.782
0.524
+0.259 [0.10, 0.40]
+0.61
0.011
0.065
Llama-8B
L3
0.449
0.269
+0.180 [0.07, 0.29]
+0.27
0.009
0.065
Llama-8B
L4
0.635
0.338
+0.297 [0.13, 0.45]
+0.55
0.007
0.059
Gemma-4B
L0
0.509
0.595
-0.086 [-0.22, 0.04]
-0.17
0.146
0.585
Gemma-4B
L2
0.658
0.583
+0.075 [-0.08, 0.22]
+0.16
0.394
1.000
Table 17
Property
Measurement
Source
Persona adapter hot-swap (Goku → Makima)
1.70 ms median, 1.41–2.15 ms (8 warm runs); 53.9 ms first call
Turn 2-A logs
Logic path invocations on persona switch (cache hit)
0 in 9/9 runs; JSON hash match 9/9
Turn 2-A logs
How-only latency on cache hit
7.1 s median (5.9–13.1 s)
Turn 2-A logs
Cache invalidation on vitals change / user override
8/8 and 2/2 misses, What re-run, constraints updated
Turn 2-B, 2-C logs
Unified memory, base + adapters, consecutive swaps
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply. We present Persona-Execution Separation (PES): persona and execution reside in different trust domains, connected by a governed contract bridge. The persona is singly-homed and may drift; execution is faceless and audited. Status summaries may return; data bodies remain in the restrictive domain except a data-loss-prevention (DLP) exception; identity stays continuous. An approval matrix, DLP, and audit enforce the crossing. PES follows from three goals: free drift, execution traceability, and decoupling. Under LLM representational indistinguishability, any single-domain mechanism meeting all three must re-introduce typed change objects, an external gate, and a stable audit anchor: PES rebuilt at higher coupling cost. A development/pilot case in a regulated platform records five decisions over one month, four with rejected alternatives. A mechanism check found no execution-side re-validation under persona perturbation (five configurations) and no persona fingerprint on hard-asserted fields of completed runs. A controlled replication in regulated coding agents reproduced the separation under isolation across five models and four providers; bridge overhead was under 0.2% of end-to-end time in both environments. A probe of a pre-separation build found the execution path decoupled from the persona by omission, not by construction. The pattern applies when multi-user deployment, execution audit, and persona churn hold jointly.
Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can continue beyond finite windows. That is safe only when dropped information is no longer needed or has been internalized. Plans are the stress case: they are written early, used for many steps, and first to be evicted. We introduce replay pairing, a diagnostic that runs the same trajectory with and without the plan in history and measures hidden-state cosine distance. On Llama-3.1-70B, plan signal spikes to 0.453 one step after the plan, then falls 4.1x in a single action-observation step; HotpotQA falls 12.4x. This is evidence that standard LLM agents do not carry plans forward as persistent state, and instead depend on the plan remaining in context. A layer-L32 probe detects this decay as a diagnostic, not as proof that it reads plan content itself. Reasoning models add a measurement confound: their <think> traces re-derive plan content, so standard stripping leaves plan evidence in the stripped condition. We name this the reasoning-trace confound and fix it with strict stripping, which removes prior <think> blocks from the stripped run only. It recovers +163% of the step+1 signal in-sample and +153% held out, while not meaningfully changing non-reasoning Llama (+4.8%). On DeepSeek-R1-Distill-Llama-70B, a Llama-trained probe transfers at AUROC 0.748 (p=6e-4), while R1-specific probes reach 1.000, suggesting R1 encodes plan signal in a different hidden-state direction. Finally, a compression stress test shows the practical cost: naive plan eviction cuts ALFWorld success by 34.7pp, while probe-gated re-surfacing does not recover it. The contribution is a measurement and stress-test framework showing that agent-critical information can be context-resident rather than persistent. Context management is load bearing, but plan protection alone is not enough.
A frontier language model's acknowledged "helpful programming assistant" persona does not survive long agentic-coding sessions in the deployment regime that production products actually run. After hours of tool-using debugging, a model that initially hedges preferences ("I don't have preferences") may begin asserting them ("Python - the feedback loop is instant..."), revealing user-visible drift that deployer evaluations may miss. Existing persona-stability studies focus on short dialogues and report little shift, leaving real-world code-generation regimes - thousands of tool-using turns, compaction, and hours-long sessions - largely uncharacterized. We introduce ContextEcho, a benchmark and reusable harness for measuring persona drift at deployment scale. It combines a 25-probe identity suite, a snapshot-then-probe protocol that forks conversation state without perturbing the main session, complementary judged and judge-free measurement surfaces, and three anonymized Claude Code sessions spanning 3,746-9,716 turns. Across 23 frontier models, ContextEcho shows that persona drift is general across organizations rather than family-specific, that in-session compaction does not reliably reset it, and that a single-shot anchor restores the trained register across measured targets. It also reveals mode-dependent downstream effects: while drift can facilitate tool-using continuation, in tool-free chat it breaks formatting contracts and inflates output length. Overall, ContextEcho provides researchers and deployers an open-source framework to audit whether the persona a model ships with is the persona users encounter at session end, across chat-completions API targets and without retraining.