cs.LGDec 24, 2025

Meta-RL with Bayesian Linear Task Models

Authors: Jingyang YouHanna Kurniawati

Abstract

Deep Bayesian reinforcement learning adapts to unseen tasks by inferring latent transition and reward models, but existing methods typically rely on variational posteriors and evidence lower bounds, introducing approximation error and unstable task representations. We introduce GLiBRL, a deep Bayesian RL framework that combines generalised linear task models with learnable non-linear basis functions. GLiBRL features conjugate Bayesian inference, yielding exact, sequential posterior updates over task parameters and model noise, together with a closed-form marginal likelihood that eliminates variational inference. The update is naturally permutation-invariant, allowing GLiBRL to integrate with both off- and on-policy algorithms. GLiBRL also learns task representation admitting an exact kernel identity, relating distances between task representations to kernel discrepancies over the task contexts. Compared against eight representative or recent meta reinforcement learning methods, GLiBRL achieves the highest aggregate zero-shot test performance on both the MuJoCo locomotion and MetaWorld manipulation benchmarks.

Explore similar work

Mar 3, 2026cs.LG

Temporal Consistency Improves Generalization in Contextual Offline Meta Reinforcement Learning

Offline meta-reinforcement learning seeks to learn a policy that generalizes to new related tasks online. Context-based methods infer a task representation from transition histories, yet learning an effective task representation without supervision remains challenging. Existing methods relying on contrastive learning learn discriminative task representations, but fail to identify task-specific dynamics, while relying on reconstruction can be insufficient to model long-horizon dependencies, limiting generalization to new tasks. We investigate the impact of temporal consistency in latent space on task representation learning, showing that enforcing multi-step predictions in latent space encourages task representations that are able to capture task-dependent dynamics while preventing representation collapse. We provide theoretical analysis characterizing sources of error in value estimation and show through extensive experiments on MuJoCo, Contextual DeepMind Control, and MetaWorld benchmarks that temporal consistency significantly improves both zero-shot and few-shot generalization.
Mohammadreza Nakheai, Aidan Scannell, Kevin Luck +1
Jul 2, 2026stat.ML

Full Bayesian Reinforcement Learning via LF-IBIS

Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards. Among RL methods, Bayesian Reinforcement Learning (BRL) addresses common practical challenges related to data scarcity by leveraging prior knowledge about the environment and sequential belief updates. However, most BRL approaches require an explicit likelihood function, which is frequently inaccessible or intractable in real-world settings. We propose Likelihood-Free Iterated Batch Importance Sampling (LF-IBIS), a novel algorithm for BRL that updates the agent's beliefs online as new interactions become available. By combining Approximate Bayesian Computation with Iterated Batch Importance Sampling, LF-IBIS enables full Bayesian inference in settings where the environment dynamics are not described by an explicit or tractable likelihood. The method yields approximate posterior distributions over both environment parameters and optimal policies, providing a quantification of policy uncertainty useful for a Bayesian treatment of the exploration-exploitation trade-off. We test the method on a simulation study in response-adaptive randomization in clinical trials, where closed-form posteriors enable validation. Additional experiments address settings where the posterior has no closed form and illustrate online policy updating based on the posterior distribution of the optimal policy.
Stefano Masini, Cecilia Viscardi, Michela Baccini
May 28, 2025cs.LG

Fully Offline Reinforcement Learning

Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance. We introduce SOReL, a fully offline Bayesian model-based RL method that learns a posterior over dynamics, estimates policy value via predictive uncertainty, and enables complete offline hyperparameter selection. We further propose TOReL, which extends this tuning framework to arbitrary model-free and model-based ORL algorithms. We provide a regret analysis showing that Bayesian offline RL achieves the minimax-optimal parametric rate under standard regularity conditions. Together, our methods establish a practical and theoretically grounded framework for fully offline RL.
Mattie Fellows, Clarisse Wibault, Uljad Berdica +3