cs.LGSep 25, 2026

Trust Guided Decision Transformer

Authors: Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath

Organizations: International Institute of Information Technology Bangalore · IBM Research Bangalore

Abstract

Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer

    May 7, 2026Yongyi Wang, Hanyu Liu, Lingfeng Li +6Transformer ArchitecturesSequence Modeling

  2. Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement

    May 7, 2026Haodong Liang, Lifeng LaiOffline Reinforcement LearningIn-Context Learning

  3. Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning

    Sep 20, 2026Suman Banerjee, Hiroyasu TsukamotoOffline Reinforcement LearningValue Functions