cs.LGMay 1, 2026

Interpretable experiential learning based on state history and global feedback

Authors: Anton Kolonin

Organizations: 1The Artificial Intelligence Research Center, Novosibirsk State University, Russia · 2Aigents, Russia.

Abstract

A new interpretable experiential learning model based on state history and global feedback is presented. It is capable of learning a behavioral model represented by a transition graph between sets of states, with transitions attributed with utility and evidence count. This model is expected to be suitable for solving reinforcement learning problem in resource-constrained environments. The model was thoroughly evaluated on the OpenAI Gym Atari Breakout benchmark, demonstrating performance comparable to some known neural network-based solutions.

Explore similar work

Jul 20, 2026cs.LG

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.
Tianzhu Ye, Li Dong, Guanheng Chen +4
Jun 19, 2026cs.LG

Inverting the Bellman Equation: From Q-Values to World Models

Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: instead of learning a model of the transition kernel PP, model-free agents typically estimate value functions tied to a specific policy and reward. In this paper, we challenge this dichotomy by proving that value-based agents trained on a sufficiently rich set of reward functions, e.g. using goal-conditioned RL, implicitly encode a unique and accurate world model. To extract this model in practice, we introduce \textit{PP-learning}, an inverse analogue to QQ-learning that samples from an agent's QQ-values, policies and rewards to decode its internal model of the environment. We then provide sufficient conditions on the type and number of goals for which agents encode the true kernel PP, covering both stochastic and deterministic MDPs over finite or continuous state spaces. Even when our assumptions are violated, we empirically demonstrate that agents trained on a handful of reward functions encode accurate dynamics in Reacher\texttt{Reacher}, MountainCar\texttt{MountainCar} and stochastic variants of FourRooms\texttt{FourRooms}. Surprisingly, we find that policies trained exclusively on a \texttt{Reacher} agent's implicit world model are quasi-optimal on out-of-distribution, velocity-based goals despite position-only training -- suggesting that agents contain hidden generalisation capabilities and providing a new lens into the connection between model-based, model-free, and goal-conditioned RL.
Alistair Letcher, Mattie Fellows, Alexander D. Goldie +3
Jul 11, 2026cs.LG

When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal

Does a reinforcement-learning agent that earns reward learn its task's hidden state? We study this question with hidden finite automata that the agent partially controls. Because each automaton is known, we can normalize reward by the best achievable return and probe the network for the true state at every step. Together the two measurements separate failures that reward alone conflates. An agent can encode too little of a state its network could hold, or encode the state and still control poorly. Weak on-policy RL matches random play while the state probe stays at chance. State learning depends on the optimizer, the training budget, and the task's structure. Permutation automata provide a warning before training: no input symbol maps two distinct states to the same successor. On a stratified held-out set, 86 of 103 permutation automata fail the state probe, and the classification is stable across probe read-outs and recovery thresholds. Most of these failures come with weak reward. High reward without the state occurs but is rare. Non-permutation automata can also fail. Oracle-normalized reward alone therefore does not establish that the task's state was learned.
James E. Allchin