What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert's behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use. We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances. A sole-carrier measurement protocol separates behavioral sufficiency, excess code rate, and information carried by observations or other memory paths; experimental bit requirements refer to the induced symbolic behavioral model under the stated occupancy. Across manipulation tasks, learned code rates remain near zero- and two-bit requirements as hidden modes grow to 512, and anticipatory memory follows a 2→1→0 requirement despite zero instantaneous demand during waiting. Learning this representation remains difficult: event-agnostic future-behavior supervision yields 36/40 sufficient seeds with one frozen configuration and improves the longest-horizon pixel setting from 0/8 to 6/8 sufficient held-out seeds (closed-loop success from 0.08 to 0.57). On unmodified community benchmarks, the protocol certifies delay-independent requirements, which sufficient codes match at mid-delay. The supervision aids commitment but can induce predictive surplus; annealing it lets imitation and rate training reduce that surplus, separating the information-theoretic target from the ability to learn it.
Figures & tables
training target
gap 6
gap 10
gap 20
M=16
M=32
teacher internal state
2/8
0/8
0/8
0/8
1/8
task-informed future behavior
8/8
8/8
8/8
8/8
5/8
event-agnostic future behavior
8/8
7/8
8/8
6/8
7/8
event-agnostic: closed-loop success
.99
.87
.90
.79
.93
Table 1: Learning near the certified two-bit requirement. First-gap rates are reported below; counts give full-gate sufficient seeds out of eight, and the last row gives all-seed closed-loop success. The event-agnostic configuration is frozen; all 40 verdicts reproduce on independent recordings. The task-informed reference uses event structure.
learner
first → second gap (bits)
sufficient
closed-loop success
plain
1.99→1.52
1/16
0.14
event-agnostic, annealed
2.00→1.13
13/16
0.65
task-informed reference
2.00→1.00
15/16
0.61
Table 2: Pixel code rates versus the certified requirement. A ′ at gap 20 requires 2→1 bits. Sixteen seeds per row; rates average sufficient seeds, success all seeds. Inputs retain probe and phase channels. The 60% schedule was selected on seeds 0–7 and tested unchanged on seeds 8–15.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
finite toys
signpost corridor
A ′ (distortion-constrained)
A ′ (attainability)
Task A
domain
finite POMDPs
grid navigation
Franka, kinematic attach
Franka, kinematic attach
Franka, physics-simulated 2 kg grasp
regime (solver)
exact; (A2), (A4), transitive
exact at W=1 ; certified at W=3,5,…,13
exact
exact
exact: ∼t transitive, (A4) fails at one step
learner
raw VQ, lr 3×10−4
raw VQ, tanh-bounded state
plain − R
scaffold-RF
plain − R
β
0.03
0.01
≤10−3
2×10−3
10−3
data, steps
6k
6k
N=1152/3584 ; 10k
N≈4 k
N=448 ; 10k
role
exact validation
cross-domain replication
distortion-constrained
attainability only
matched performance
Appendix
Table 3: Benchmark families, regimes and frozen primary configurations ( N = training episodes; K=16 codes throughout; the Franka runs use batch 512).
choice
premise
evidence
Ct is the only carrier across time
Definition 2
Appendix B.2 : bypass under-counts in three domains
no empirical advantage over an unconditional prior in these families
forecaster used at training only
sole carrier at evaluation
forecaster removed before every reported rate and rollout; Ct is the only cross-time variable (§ 5 )
Appendix
Table 4: Design choices, premises and evidence (§ 4 ).
phase
β1 (0.50)
β2 (0.50)
θ (0.03)
mass bin (0.25)
distractor z (0.25)
identification (scan)
1.00
1.00
0.10
0.26
—
gap 1
0.53
0.51
0.04
0.23
—
grasp
0.99
0.47
0.05
0.24
—
gap 2
0.51
0.46
0.03
0.18
0.99
place
0.49
1.00
0.04
0.26
—
Appendix
Table 5: Leak probe on A ′ (tier 4, gap 6): held-out accuracy ( n=128 episodes) of a classifier on raw observation windows, per phase and hidden variable; chance in parentheses. Bold entries are the steps at which the visibility convention marks the class visible.
gap
architecture
β
sufficient
first-gap rate
post-use rate
behav. error
closed-loop success
6
hierarchical
0
8/8
2.03±0.08
1.00
0.000
1.00
6
hierarchical
10−3
8/8
2.09±0.17
1.00
0.000
1.00
6
intent
0
8/8
2.02±0.05
1.09
–
0.96
6
intent
10−3
8/8
2.03±0.08
1.00
–
0.94
20
hierarchical
0
8/8
2.04±0.10
1.00
0.027
0.96
20
hierarchical
10−3
8/8
2.03±0.08
1.00
0.000
1.00
Appendix
Table 6: Strict read-side architectures on A ′ with behavioral-forecast supervision (8 seeds per cell; theory 2.00→1.00 ; forecaster removed at evaluation). Without the forecast target the same architectures are sufficient on 0/8 seeds in every cell.
δ (mm)
0
0.5
1
2
4
8
closed-loop success
0.59
0.52
0.61
0.83
0.84
0.96
slot accuracy
0.68
0.68
0.79
0.93
0.95
0.98
held-out action error (MSE)
0.107
0.116
0.053
0.015
0.022
0.007
behaviorally correct seeds
5
5
7
15
14
16
seeds passing the sufficiency gate
5
4
7
4
6
3
flag A: correct, rate <0.9 bit
0/5
0/5
0/7
0/15
1/14
0/16
Appendix
Table 7: Injected-leak audit. Top: 16 seeds per magnitude δ . Correct means held-out class error below 0.05 at the first place step. Flag A, fixed before the runs, pairs correctness with second-gap rate below 0.9 bit. The exploratory flags pair correctness with failure of the full gate (B) or second-gap memory alone (B ′ ). Bottom: false alarms on separate clean recordings, conditioned on correctness.
CPU, batch 1 (ms)
GPU, batch 64 (ms)
T
Transformer
GRU
DIACRITIC
Transformer
GRU
DIACRITIC
33
1.13
0.55
0.54
1.77
1.12
1.18
128
1.78
0.55
0.55
1.83
1.18
1.17
256
3.51
0.56
0.53
2.12
1.15
1.18
512
9.03
0.55
0.54
5.45
1.10
1.13
1024
29.3
0.56
0.55
16.8
1.13
1.17
Appendix
Table 8: Per-step inference latency (median of 200 steps) as a function of the history length T , with the widths of the paper’s models. The Transformer re-encodes its token history at every step (no key–value cache), so these timings describe that implementation; its carried history is 64(T−1) floats (8.2 kB at T=33 , 262 kB at T=1024 ), compared with 128 bytes for the GRU and 129 bytes for DIACRITIC (state plus an 8-bit code index).
parameters
training time
at deployment: carried state / per-step latency
forecaster, 3000 steps (default)
0.49 – 0.51 M
11 – 12 s
removed
random-offset forecaster, 100 steps (timing only)
0.51 M
≈1 s
removed
compact student, 10 000 steps, with or without the auxiliary head
0.19 M
≈800 s
129 bytes / 0.55 ms
full-history Transformer baseline, 10 000 steps
0.49 M
≈650 s
64(T−1) floats / 1.1 – 29 ms for T=33 – 1024
Appendix
Table 9: Cost bookkeeping on A ′ gap 20 (one B200 GPU; medians over runs). The forecaster and the auxiliary head exist only during training.
M=4
M=8
M=16
M=32
sufficient seeds / closed-loop success
ours, K=16
8/8 / 0.91
8/8 / 0.96
8/8 / 0.98
8/8 / 0.98
ours, K=64
8/8 / 1.00
8/8 / 1.00
8/8 / 0.99
8/8 / 1.00
sys-ID, K=16
1/8 / 0.11
0/8 / 0.08
0/8 / 0.06
0/8 / 0.07
sys-ID, K=64
8/8 / 0.71
3/8 / 0.18
0/8 / 0.08
0/8 / 0.09
sys-ID, K=128
8/8 / 0.79
8/8 / 0.34
0/8 / 0.12
0/8 / 0.16
Appendix
Table 10: Task A controls, 8 seeds per cell. Top block: seeds sufficient at the place step / closed-loop success (128 episodes per seed). Bottom block: code rate during grasp (bits) / I(C;mass∣Oˉ) (bits).
M
objective
sufficient (place)
grasp rate (bits)
I(C;mass∣Oˉ)
success
128
ours, K=16 , β=10−3
8/8
0.03
0.01
0.99
128
ours, K=16 , β=0
8/8
0.18
0.02
0.99
128
sys-ID, K=16
0/8
3.73
0.62
0.04
128
sys-ID, K=256
3/8
6.84
1.40
0.17
512
ours, K=16 , β=10−3
8/8
0.03
0.01
0.99
512
ours, K=16 , β=0
8/8
0.10
0.02
0.98
Appendix
Table 11: Task A at M=128 and 512 (8 seeds per row; 128 closed-loop episodes per seed).
task
learner
sufficient
first-gap rate
transport rate
closed loop
readout-2
− R
8/8, 7/8, 8/8
2.00, 2.02, 2.03 [2.00]
1.10, 1.23, 1.06 [1.00]
0.94, 0.85, 0.87
readout-2
forecast
8/8, 8/8, 8/8
2.02, 2.08, 2.03 [2.00]
1.11, 1.27, 1.09 [1.00]
0.95, 0.90, 0.96
readout-2
sys-ID, largest K
4/8, 6/8, 1/8
4.82, 5.56, 7.49
4.49, 4.74, 7.93
0.25, 0.26, 0.19
readout-3
forecast
6/8, 5/8, 4/8
3.03, 3.09, 3.17 [3.00]
3.06, 3.01, 3.09 [3.00]
0.25, 0.28, 0.15
readout-3
forecast, K=8
0/8, 0/8, 0/8
2.73, 2.92, 2.87
2.49, 2.46, 2.76
0.05, 0.05, 0.05
readout-3
sys-ID, largest K
6/8, 0/8, 0/8
4.83, 5.78, 7.94
4.78, 5.78, 8.31
0.32, 0.07, 0.03
Appendix
Table 12: Readout task, 8 seeds per cell, K=16 unless noted (sys-ID largest K : 256, 256, 1024); each cell lists M=32 / 128 / 512 . Rates are all-seed means except for the K=16 forecast row on readout-3, which averages sufficient seeds; closed-loop success is always averaged over all seeds. Behavioral-learner verdicts replicate on the fresh recordings; the largest- K readout-2 sys-ID count at M=512 changes from 1/8 to 0/8 . Theory is in brackets.
configuration
H(C∣Oˉ)
I(C;mass∣Oˉ)
H(Cid∣Oˉ)
I(Cid;θ∣Oˉ)
success
M=32 , N=448 (every mass in training)
DIACRITIC ( K=16 )
0.34
0.01
–
–
0.99, 0.88
shared, weight 1 ( K=16 )
2.83
0.73
–
–
0.07, 0.05
shared, weight 0.1 ( K=256 )
4.52
0.80
–
–
0.38, 0.23
dual carrier ( Kid=256 )
0.46
0.03
4.12
2.73
0.90, 0.84
dual, stop-gradient
0.29
0.00
4.35
2.45
0.93, 0.88
Appendix
Table 13: Task A, sys-ID controls (8 seeds): transport rates of the behavioural code C and of the identification code Cid (bits), mass information carried by C , θ information carried by Cid , and closed-loop success (mean, min). “shared” = the θ head reads C itself.
sufficient seeds
gap 6
gap 10
gap 20
M=16
M=32
EA, anneal 60% (frozen)
8/8 ( 8/8 )
7/8 ( 7/8 )
8/8 ( 8/8 )
6/8 ( 6/8 )
7/8 ( 7/8 )
EA, anneal 40%
—
—
7/8 ( 7/8 )
—
7/8 ( 7/8 )
EA, not annealed
6/8 ( 6/8 )
8/8 ( 7/8 )
7/8 ( 7/8 )
8/8 ( 8/8 )
6/8 ( 6/8 )
TI
8/8 (—)
8/8 ( 8/8 )
8/8 ( 8/8 )
8/8 ( 8/8 )
5/8 ( 5/8 )
teacher state
2/8
0/8
0/8
0/8
1/8
rate / closed loop
gap 6
gap 10
gap 20
M=16
M=32
Appendix
Table 14: Forecast supervision on A ′ (state observations), 8 seeds per cell. Top: sufficient seeds under the full-trajectory gate (in parentheses: the same models re-evaluated on an independent recording). Bottom: first-gap → second-gap rate of the sufficient seeds (bits; theory 2.00→1.00 ) / closed-loop success over all seeds (128 episodes per policy). EA = event-agnostic random-offset forecast; TI = task-informed forecast.
horizon
learner
sufficient
rate (bits)
success
slot accuracy
gap 6
plain − R
2/8
1.99→1.00 ( n=2 )
0.17
0.60
gap 6
task-informed forecast
7/8
1.99→1.12
0.25
0.94
gap 20
plain − R
1/16
1.99→1.52 ( n=1 )
0.14
0.56
gap 20
event-agnostic, not annealed
7/8
2.15→2.01
0.40
0.80
gap 20
event-agnostic, annealed from 40%
12/16
2.01→1.13
0.51
0.82
gap 20
event-agnostic, annealed from 60%
13/16
2.00→1.13
0.65
0.89
Appendix
Table 15: Pixel A ′ : sufficient seeds (full-trajectory gate), rates of the sufficient seeds (theory 2.00→1.00 ), and closed loop over all seeds (128 episodes per policy).
target
gap 6
gap 10
gap 20
M=16
M=32
task-informed forecast of pending behavior
8/8
8/8
8/8
8/8
5/8
teacher’s internal state
2/8
0/8
0/8
0/8
1/8
unsupervised (best variant)
3/8
2/8
1/8
0/8
0/8
forecast: rate before use (theory 2.00)
2.01
2.05
2.09
2.08
2.13
forecast: rate after use (theory 1.00)
1.07
1.06
1.10
1.09
1.11
forecast: closed-loop success
0.95
0.89
0.92
0.84
0.73
Appendix
Table 16: Distillation controls on A ′ and P-I (8 seeds per cell; full-trajectory gate of § 4.1 ; the event-agnostic objective is in Table 14 ). Top block: seeds sufficient. Bottom block: rates and closed-loop success of the sufficient seeds under the forecast target.
variant
sufficient
post-use rate (bits)
rate measurement
success
DIACRITIC ( − R, β=10−3 )
✓
1.38 (near)
✓
0.97
scaffold → DIACRITIC ( N≈4 k)
✓
0.99 – 1.05 on 5/6 (strongest)
✓ ( α=0 at evaluation)
0.87 – 1.00
β=0
✓
1.53 (redundant)
✓
0.96
continuous bottleneck
possible
KL bound 6.9 ; nuisance 0.32
upper bound only
—
GRU bypass
✓
1.06 (apparent under-rate)
× (not the sole carrier)
0.97 – 1.00
system identification
budget-dep.
2.2 – 3.2 (learned rate)
✓
0.06 – 0.79
Appendix
Table 17: Sufficiency and minimality on A ′ (gap 6, N=1152 unless noted). Minimality is the post-use rate against the 1.00 -bit boundary. The variational KL bound of the continuous model is not comparable to hard-code entropy and is therefore not read as a minimality result. Appendices D.4 and C.3 report additional variants.
policies
n
SΓ (gap 2 )
slot accuracy
success
action error (cm)
teacher-forced sufficient ( − R, seeds 0, 6)
2
1.00 , 0.95
0.92 , 0.94
0.66 , 0.36
3.1 , 5.7
insufficient, − R
6
≤0.08
0.42 – 0.66
0.09 – 0.27
8 – 12
insufficient, scaffold-R
8
≤0.07
0.41 – 0.54
0.10 – 0.42
7 – 10
Appendix
Table 18: Closed-loop evaluation of the 16 pixel policies (128 episodes each; SΓ under the policy’s own occupancy).
cue bits, variant
L=5
L=10
L=20
L=40
L=80
1 bit, − R
8/8 / 1.25
8/8 / 1.19
7/8 / 1.00
3/8 / 1.00
5/8 / 1.00
1 bit, − RF
8/8 / 1.12
8/8 / 1.50
8/8 / 1.38
5/8 / 1.10
6/8 / 1.17
1 bit, scaffold-RF
8/8 / 1.17
8/8 / 1.00
8/8 / 1.06
8/8 / 1.19
6/8 / 1.08
2 bits, − R
6/8 / 2.00
4/8 / 2.00
1/8 / 2.00
0/8 / —
0/8 / —
Appendix
Table 19: T-maze: seeds sufficient / corridor rate of sufficient seeds (bits; theory log2R ).
gap 6
gap 20
learner
sufficient
closed loop
sufficient
closed loop
plain − R
1/8 / 1/8
0.37 / 0.51
0/8 / 0/8
0.37 / 0.24
task-informed forecast
3/8 / 3/8
0.21 / 0.51
2/8 / 2/8
0.21 / 0.35
task-informed forecast from step 0
8/8 / 8/8
0.10 / 0.30
7/8 / 7/8
0.07 / 0.32
event-agnostic, not annealed
2/8 / —
0.19 / 0.51
0/8 / —
0.13 / 0.35
system identification
8/8 / —
0.08 / 0.38
7/8 / —
0.06 / 0.40
Appendix
Table 20: Weighing task, 8 seeds per cell: sufficient seeds on the training recording / on the independent recording, and closed-loop success / grasp-side accuracy over all seeds (chance 0.25 ).
learner
L=10
L=20
L=50
L=100
GRU policy, continuous state
8 (1.00)
8 (1.00)
8 (1.00)
7 (0.93)
compact carrier, plain imitation
1 (0.53)
1 (0.53)
0 (0.46)
0 (0.46)
compact carrier, event-agnostic forecast
8 (1.00)
7 (0.94)
5 (0.80)
0 (0.48)
compact carrier, − RF (decoder sees future observations)
—
1 (0.53)
0 (0.46)
0 (0.46)
compact carrier, annealed continuous scaffold
—
0 (0.46)
0 (0.47)
0 (0.47)
GRU policy, 8-dimensional state
—
0 (0.47)
0 (0.41)
0 (0.35)
Appendix
Table 21: Passive T-maze: seeds (of 8) that solve the task in closed loop (success ≥0.99 ); mean success in parentheses.
task
delay
certified
suff., plain
suff., event-agnostic
learned rate
MemoryChain, 1 bit
10
1.00
7/8
8/8
1.00
MemoryChain, 1 bit
30
1.00
4/8
8/8
1.00
MemoryChain, 1 bit
100
1.00
0/8
2/8
1.00
MemoryChain, 2 bits
10
1.99
7/8
7/8
1.99
MemoryChain, 2 bits
30
1.99
2/8
3/8
1.99
MemoryChain, 2 bits
100
1.99
0/8
0/8
—
Appendix
Table 22: Certified requirement and learned rate on community benchmarks (8 seeds per cell). The certified value is the empirical requirement at mid-delay under the recorded occupancy ( 512 episodes, hence 1.99 and 2.99 ); the learned rate is H^(Ct∣Ot) at the same step over the sufficient seeds of both learners.
method
target representation
F
S
E
B
rate of the target’s fixed-demonstrator form under our solver
Table 23: Representation targets in related work (§ 6 ). Property columns: F the target reproduces a fixed demonstrator’s action distribution; S continuations are compared over stochastic reachable supports, so future observations act as decoder side information; E minimality is defined through occupancy-weighted conditional entropy; B the general case yields a non-transitive compatibility bracket. Individual properties appear in prior work; the contribution here is their combination.
Memory consolidation determines both what a learner can do now and which changes remain implementable later. We develop a finite-model synthesis of operational state abstraction and optimal control under the stability-evidence-revision (SER) framework. ``Poincaré meets Bellman'' names two complementary roles: qualitative dynamics identifies reusable action-response structure, and dynamic programming prices acquisition, retention, reuse, merging, and forgetting. Recurrence enters separately through the timing and value of future demands. We distinguish active quotient merging from historical information erasure, characterize exact repair by zero-error functional coding and causal migration, and derive a Bellman recursion over the joint law of hidden state and complete deployed memory. A first-return model yields an explicit retention rule. Conditional results show how factor sharing avoids enumerating combinations and how independent informative observations improve identification, while leaving some zero-error evidence budgets unchanged. Finite enumerations verify the coding and retention calculations. The synthesis gives an exact benchmark for specified finite models, without claiming universal recurrence, bounded-memory open-ended learning, or tractable global planning.
Xin Li
Department of Computer Science University at Albany
Memoir combines per-sample fast memory, shared slow parameters, variable-depth latent recurrence, and a future-latent energy objective. We test its riskiest coupling: each pondering iteration may rewrite the fast tier that the same iteration reads. On procedural associative recall with key interference, we compare a coupled arm against an otherwise identical read-only pondering arm. Both arms contain 81,738 parameters, including 76,362 trainable parameters, and use matched declared forward multiply-accumulate counts, data, optimizer, schedule, and seeds. After 240 training steps across 12 seeds, coupled recall is 0.5203 with a 95 percent interval of [0.4522, 0.5883], while read-only recall is 0.6557 with [0.5953, 0.7160]. The arms are paired per seed, and the read-only lead of 0.1354 gives a paired t of 3.23 on 11 degrees of freedom with a 95 percent interval of [0.0431, 0.2277] on the difference, winning on 10 of 12 seeds. After 960 steps across 8 seeds, both arms reach 1.0000, so the measured effect is a learning-speed penalty at a fixed budget, not a demonstrated capability penalty. That longer control is ceiling limited, leaving convergence on a non-saturating task unmeasured. A predicted failure in which memory rewriting corrupts the energy signal did not occur: the energy margin grew and held. Kernel restructuring also reduced delta-rule forward time from 0.907 ms to 0.351 ms on the stated device. Code and evidence are available at https://github.com/RightNow-AI/Memoir
Behavior cloning provides strong imitation learning guarantees when training and test environments share the same dynamics. However, in many deployment settings the test environment's transitions differ from training, and classical offline IL offers no recourse: the learner must commit to an action at every state, even when its demonstrations are uninformative and could lead to arbitrary degradation of performance. This motivates the study of selective imitation, where the learner may choose to stop when it cannot act reliably. We introduce a model for selective imitation under arbitrary dynamics shift: given labeled expert demonstrations from a training environment and unlabeled state trajectories from the same expert in a test environment, the learner outputs a selective policy that is complete (rarely stops in training) and sound (incurs low regret before stopping in test). Our algorithm, SeqRejectron, constructs a stopping rule using a small set of validator policies whose size is independent of the horizon or policy class. For deterministic policies, this yields horizon-free O~(log∣Π∣/ε2) sample complexity, assuming sparse costs. For stochastic policies, we obtain analogous horizon-free guarantees using a cumulative Hellinger stopping time. We extend the framework to misspecified experts and different expert policies across train and test and obtain results that gracefully degrade with the amount of misspecification.