Model-Based RL

RL: Reinforcement Learning

Momentum

24 papers in the last four weeks, up 243% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 142

All topics
CardsList
  1. RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation

    Oct 6, 2026Jing Xie, Shouwei Ruan, Yubin Wang +4Robot Policy LearningWorld Model Learning

  2. Considering Context: When World Models Need Context Encoders

    Oct 5, 2026Oleg Smirnov, Sofiane Ennadir, John Pertoft +2Reinforcement LearningWorld Model Learning

  3. Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

    Oct 5, 2026Shengtao Wen, Xiang Chen, Yu Tian +3World Model-Based PlanningAction-Conditioned World Models

  4. Optimal Control with Learned Critics under Unmodeled State Dependencies

    Oct 4, 2026Philipp Schoch, Markus RyllResidual Dynamics LearningRobot Dynamics Learning

  5. R2R^2-WAM: Repair-and-Reject Post-Training for World Action Models

    Oct 4, 2026Ruiyan Xu, Haisheng Su, Sixu Lin +4World Action ModelsAction-Conditioned World Models

  6. In CEM, a World Model Is Also a Proposal Mechanism

    Oct 1, 2026Oliver Obst, Frieder StolzenburgBlack-Box OptimizationModel-Based Planning

  7. Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust

    Sep 29, 2026Ziqi Wen, Ting Xu, Lianyu Wang +5Uncertainty QuantificationWorld Model-Based Planning

  8. Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction

    Sep 28, 2026Victor Paredes, Ayonga HereidHumanoid Robot LocomotionControl Barrier Functions

  9. MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

    Sep 27, 2026Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de WittGame-Playing AgentsWorld Model Learning

  10. Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC

    Sep 26, 2026Yi Xian Goh, Sze Jue Yang, Hao LuanData-Driven ControlModel Predictive Control

  11. Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality

    Sep 26, 2026Linhao Wang, Yiyan Fan, Dongjin HuangWorld Model-Based PlanningModel-Based Planning

  12. Dual-Frontier: When Can an Agent Trust Its World Model?

    Sep 22, 2026Huatai Zhu, Qiang Chen, Ziqian Kou +5World Model-Based PlanningModel-Based Planning

  13. Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning

    Sep 21, 2026Kejia Hu, Wentong Zhai, Bo Zhao +1Robotic ManipulationVision-Language-Action Models

  14. Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

    Sep 17, 2026Jonathan Hau, Alessandro AbateBayesian RLSafe RL

  15. A Convergence Framework for Deep VV-Learning: Error Propagation and Sharp Action-Gap Bounds

    Sep 16, 2026Yury KolomeytsevReinforcement LearningPolicy Learning

  16. Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning

    Sep 16, 2026Everest Yang, Skye Thompson, George D. KonidarisNon-Stationary RLContinual Robot Learning

  17. ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation

    Sep 15, 2026Yushan Bai, Boyu Zheng, Zhiyang Mao +4Robot Policy LearningContact-Rich Robotic Manipulation

  18. Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models

    Sep 12, 2026Andreas Schwung, Steve Yuwono, Sofiene Lassoued +1ManufacturingRL Control

  19. Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

    Sep 10, 2026Fernando Palafox, David Fridovich-KeilTest-Time AdaptationWorld Model Learning

  20. Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

    Sep 10, 2026Corentin Pla, Hugo Richard, Marc Abeille +1Reinforcement LearningApproximation Algorithms

  21. HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

    Sep 9, 2026Zengjue Chen, Peidong Liu, Jiawei Li +1Robot Policy LearningWorld Models

  22. CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning

    Sep 8, 2026Pietro Noah Crestaz, Mohamed Yassine Kabouri, Nicolas Mansard +1RL for RoboticsReinforcement Learning

  23. World-Model-Augmented Visual Locomotion for Humanoids on Foothold-Constrained Terrain

    Sep 2, 2026Yuxi Liu, Lijun Han, Ziming Wang +3Humanoid Robot LocomotionTerrain-Aware Robot Locomotion

  24. NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

    Sep 1, 2026Tomáš Holeček, Viliam LisýImperfect-Information GamesZero-Sum Games

  25. Reinforced Planning with Latent World Models

    Aug 19, 2026Armin Sommer, Jannik SchillingWorld Model-Based PlanningRobotic RL

  26. Scaling Automatic Research Agents via World Models

    Aug 12, 2026Xiyuan Yang, Sheikh Sarwar, Jingru Cheng +8World Model LearningRL Post-Training

  27. Dynamics Models for Offline Hyperparameter Selection in Real-World RL

    Aug 11, 2026Jordan Coblin, Han Wang, Martha White +1RL ControlOffline RL

  28. IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

    Aug 11, 2026Zefeng Liang, Jie Qiao, Ruichu Cai +2Reinforcement LearningRobust RL

  29. LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

    Aug 7, 2026Cheng Guo, Mingzhe Ni, Angelo Cangelosi +1Humanoid Loco-ManipulationLatent Action Learning

  30. Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection

    Aug 7, 2026Jiazhuo Li, Yiming Fei, Zhiruo Zhou +1Action-Conditioned World ModelsLatent World Models

  31. Analytic Planning under Uncertainty with Moment Closure

    Aug 3, 2026Shishir Sharma, Doina PrecupDecision-Making under UncertaintyModel-Based RL

  32. Climate-Dyna Deep Hedging for XVAs: Model-Based Reinforcement Learning, Residual Climate HVA, and Hedge-Instrument Discovery

    Aug 2, 2026Xiaozhen Wang, Francois Buet-GolfouseQuantitative FinanceResidual Policy Learning

  33. DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

    Jul 31, 2026Jiayang Niu, Yan Wang, Jie Li +4World ModelsQuantum RL

  34. Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

    Jul 31, 2026Anders Jonsson, Emilie Kaufmann, Gianmarco Tedeschi +1Hierarchical RLPolicy Learning

  35. QQWorld: Quantile-Quantile Matching for World Model Regularization

    Jul 30, 2026Zhoushun Yu, Xiaoyu Hu, Xiangyu XuWorld Model LearningWorld Model-Based Planning

  36. Learning Implicit Causal World Models from Multi-Agent Demonstrations

    Jul 28, 2026Jasorsi GhoshCausal Representation LearningCausal World Models

  37. Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance

    Jul 28, 2026Gaspard Lambrechts, Adrien Bolland, Daniel Ebi +1Reinforcement LearningWorld Models

  38. Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

    Jul 22, 2026Taisuke Takayama, Naoto Yoshida, Tadahiro TaniguchiPredictive CodingPartially Observable RL

  39. Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination

    Jul 22, 2026Jiaqi Li, Xinglong Zhang, Haibin Xie +3Koopman Operator LearningLatent Dynamics Modeling

  40. Reinforcement Learning: From Algorithms To Foundation Models

    Jul 20, 2026Zihan DingReinforcement LearningRL for Video Generation

  41. Enhancing Personalized Bladder Cancer Treatment Through Reinforcement Learning: A Recurrent Patient State Transition Decision Support Framework

    Jul 18, 2026Divyansh Chawla, Anshu Garg, Isshaan SinghReinforcement LearningHealthcare

  42. Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL

    Jul 18, 2026Zhaohui WangRisk-Sensitive RLModel Predictive Control

  43. Certifiable Safe Model-Based Reinforcement Learning with Control-Affine Dynamics Approximation

    Jul 17, 2026Hao Zhou, Yanze Zhang, Cameron Reid +1Safe RLModel-Based RL

  44. RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences

    Jul 15, 2026Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian RamamoorthyWorld Model LearningOffline RL

  45. Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

    Jul 14, 2026Ilias Kazantzidis, Timothy J. Norman, Yali Du +1Safe RLReward Design for RL

  46. Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

    Jul 10, 2026Guanquan Wang, Yoshimasa TsuruokaTrajectory GenerationOffline RL

  47. Solving Markov Decision Processes with Future Information via MPC

    Jun 23, 2026Shambhuraj Sawant, Akhil S Anand, Dirk Reinhardt +1RL ControlMarkov Decision Processes

  48. Inverting the Bellman Equation: From QQ-Values to World Models

    Jun 19, 2026Alistair Letcher, Mattie Fellows, Alexander D. Goldie +3Reinforcement LearningModel-Free RL