cs.AISep 27, 2026

Agentic Multi-Turn Reasoning: A Fairness Approach

Authors: Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren, Marios Savvides, Khoa Luu

Organizations: CVIU Lab, University of Arkansas, USA · Dep. of Physics, University of Arkansas, USA · Dep. of Geosciences, University of Arkansas, USA · Carnegie Mellon University, USA

Abstract

Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or ΦΦ-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Adaptive Latent Agentic Reasoning

    Jun 1, 2026Dongwon Jung, Peng Shi, Yi Zhang +2Agentic ReasoningEfficient Latent Reasoning

  2. AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

    Aug 30, 2026Xinke Jiang, Yue Fang, Zhibang Yang +12Agentic Reinforcement LearningAgentic Benchmarks

  3. CANTANTE: Optimizing Agentic Systems via Contrastive Credit Attribution

    May 13, 2026Tom ZehlePrompt OptimizationScoring