cs.LGNov 28, 2025

Poincaré Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments

Authors: Xin Li

Organizations: Department of Computer Science University at Albany

Abstract

Memory consolidation determines both what a learner can do now and which changes remain implementable later. We develop a finite-model synthesis of operational state abstraction and optimal control under the stability-evidence-revision (SER) framework. ``Poincaré meets Bellman'' names two complementary roles: qualitative dynamics identifies reusable action-response structure, and dynamic programming prices acquisition, retention, reuse, merging, and forgetting. Recurrence enters separately through the timing and value of future demands. We distinguish active quotient merging from historical information erasure, characterize exact repair by zero-error functional coding and causal migration, and derive a Bellman recursion over the joint law of hidden state and complete deployed memory. A first-return model yields an explicit retention rule. Conditional results show how factor sharing avoids enumerating combinations and how independent informative observations improve identification, while leaving some zero-error evidence budgets unchanged. Finite enumerations verify the coding and retention calculations. The synthesis gives an exact benchmark for specified finite models, without claiming universal recurrence, bounded-memory open-ended learning, or tractable global planning.

Figures & tables

Explore similar work

Sep 22, 2026cs.LG

Minimal Recurrent Behavioral Memory for Imitation under Partial Observability

What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert's behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use. We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances. A sole-carrier measurement protocol separates behavioral sufficiency, excess code rate, and information carried by observations or other memory paths; experimental bit requirements refer to the induced symbolic behavioral model under the stated occupancy. Across manipulation tasks, learned code rates remain near zero- and two-bit requirements as hidden modes grow to 512512, and anticipatory memory follows a 2→1→02\to1\to0 requirement despite zero instantaneous demand during waiting. Learning this representation remains difficult: event-agnostic future-behavior supervision yields 36/4036/40 sufficient seeds with one frozen configuration and improves the longest-horizon pixel setting from 0/80/8 to 6/86/8 sufficient held-out seeds (closed-loop success from 0.080.08 to 0.570.57). On unmodified community benchmarks, the protocol certifies delay-independent requirements, which sufficient codes match at mid-delay. The supervision aids commitment but can induce predictive surplus; annealing it lets imitation and rate training reduce that surplus, separating the information-theoretic target from the ability to learn it.
Aug 13, 2026cond-mat.stat-mech

Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value

What a finite learning device has recorded and what will hold value for it on future tasks are not the same quantity. We develop a typed accounting for finite-state learning devices that separates four components: a training-side fit functional ΦfitΦ_{\mathrm{fit}}, the record-correlation stock JD=I(M;D)J_{D}=I(M;D), an update-side search ledger σMσ_{M}, and an operational capital value V(M;T,b)V(M;T,b). This value is the work gap between an informed protocol class and a blind class obtained by deleting the memory-read port and re-optimizing from scratch. (I) Separation: for every nn, there is a device family on which record correlation and world correlation grow by nln⁡2n\ln 2 while the capital gain is exactly zero. In the flat∗\mathrm{flat}^{*} regime, data-free updates never increase VV. (II) Capitalization ledger: an exact flat∗\mathrm{flat}^{*} extraction identity and a universal ledger identity give, for (F5′')-stable MM-local updates under a no-discarded-record-correlation condition (f), the bound ηcap≤1η_{\mathrm{cap}}\le 1 for the capitalization efficiency ηcap=ΔV/(kT σM)η_{\mathrm{cap}}=ΔV/(k T\,σ_{M}), together with necessary and sufficient conditions for equality. (III) Value retention: for the retention gap LgenL_{\mathrm{gen}} and retention ratio ρgenρ_{\mathrm{gen}} (the former carries no sign constraint; the latter is defined for positive training-side value and is not confined to [0,1][0,1]) we give a two-layer alignment domain: an exact exchange rate between value and the side-information-adjusted record fit I(M′;D∣Y)I(M';D\mid Y) without any record-side-information independence assumption, and a raw record-stock exchange rate under a joint side-information neutrality condition (M,D)⊥Y(M,D)\perp Y, whose boundary is marked by an explicit one-time-pad witness. These are statements about finite-device value retention under task-distribution shift, not a theory of statistical generalization.
May 11, 2026cs.LG

Consolidation-Expansion Operator Mechanics:A Unified Framework for Adaptive Learning

Every adaptive learning system must alternate between two operations: consolidating what it already knows and expanding into new evidence. We propose \emph{Consolidation-Expansion Operator Mechanics} (OpMech), a framework that makes this structure precise. The central object is the \emph{order-gap} \Ogap(θ;e)\Ogap(θ; e), the degree to which a consolidation operator~QQ and an expansion operator~PeP_e fail to commute at a given knowledge state. Because the order-gap is computable from the system's own trajectory, it serves as a real-time control signal: large values indicate that the system is still sensitive to the ordering of consolidation and expansion; once the order-gap falls and stays small, further processing is unlikely to change the outcome. Three results give the signal precise meaning: the order-gap decays along convergent trajectories; a persistently large order-gap implies the system is far from its settled state; and an order-gap-based stopping rule terminates with provable guarantees in both noiseless and bounded-noise settings. The framework applies across five domains: bandits, reinforcement learning, stochastic optimization, continual learning, and recursive language models. We give conditions under which the order-gap reliably tracks convergence in three representative cases. We develop the recursive language model application in detail, showing how OpMech replaces heuristic stopping rules and fixed recursion budgets with principled, evidence-driven alternatives.