On-Policy

Recent momentum

-75%

3 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in On-Policy.

Period ending 2026-09-07

1 new paper

A weekly snapshot of new work published in On-Policy.

48 papers

Latest in On-Policy

Open your feed →
CardsList
  1. Continual Learning in Transition

    Aug 6, 2026Zhiyan Hou, Dan Zhang, Tao Feng +11Continual Learning MethodsTransitions

  2. Reversal Q-Learning

    Jun 16, 2026Aditya Oberai, Seohong Park, Sergey LevineFlow-Based PolicyOffline Reinforcement Learning

  3. ExpRL: Exploratory RL for LLM Mid-Training

    Jun 15, 2026Violet Xiang, Amrith Setlur, Chase Blagden +2Large Language Model ReasoningSparse Rewards

  4. Extreme Region Policy Distillation

    May 25, 2026Changyu Chen, Xiting Wang, Rui YanOn-PolicyOff-Policy Evaluation

  5. ECHO: Terminal Agents Learn World Models for Free

    May 23, 2026Vaishnavi Shrivastava, Piero Kauffmann, Ahmed Awadallah +1Agentic Reinforcement-Learning FrameworkOn-Policy

  6. TADPO: Reinforcement Learning Goes Off-road

    Mar 6, 2026Zhouchonghao Wu, Raymond Song, Vedant Mundheda +3Proximal Policy OptimizationTerrain