cs.CLSep 30, 2026

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Authors: Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao, Fang Kong

Organizations: Southern University of Science and Technology · City University of Hong Kong

Abstract

Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box2^2-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box2^2-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

    Jun 11, 2026Ziyi Wang, Yuxuan Lu, Yimeng Zhang +8AI Agent ReliabilityReinforcement Learning

  2. Learning When to Trust via Selective Context Preference Optimization

    Aug 6, 2026Xian Sun, Wei Chow, Yingshuo Wang +4Knowledge Conflicts in Language ModelsLLM Reliability

  3. Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

    Nov 14, 2025Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan LiRL for Language ModelsLLM Hallucination Mitigation