cs.LGOct 5, 2026

Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity

Authors: Lan Shi, Daigo Shishika, Xuan Wang

Organizations: George Mason University, USA

Abstract

Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.

Figures & tables

Explore similar work

CardsList
  1. Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

    Sep 30, 2026Morgan Byrd, Jacob Blevins, Maks Sorokin +2Neural PoliciesRapid Adaptation

  2. Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?

    Apr 20, 2026Ku Onoda, Paavo Parmas, Manato Yaguchi +1Policy GradientDifferentiable Physics

  3. Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning

    Nov 25, 2025Charlotte Beylier, Hannah Selder, Arthur Fleig +2Offline Reinforcement LearningSaliency