Organizations: CVIU Lab, University of Arkansas, USA · Dep. of Physics, University of Arkansas, USA · Dep. of Geosciences, University of Arkansas, USA · Carnegie Mellon University, USA
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or Φ-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.
Figures & tables
Figure 1: Overview of the proposed Φ -MPO approach. Left: Φ -MPO consistently outperforms GPT-4o, AgentFlow + FlowGRPO across four benchmarks. Right: Compared with DPO and Flow-GRPO, Φ -MPO addresses both long-horizon reasoning and imbalanced problems.
Figure 2: The Influence of Imbalance Data.
Figure 3: Our Proposed Φ -MPO Framework .
Method
Model Size
Knowledge-Intensive Search
Agentic
Bamboogle
2Wiki
HotpotQA
Musique
GAIA
Qwen2.5-7B-Inst
7B
12.0
23.0
21.0
6.0
3.2
Qwen2.5-14B-Inst
14B
21.6
26.7
20.0
8.0
5.5
Qwen2.5-32B-Inst
32B
24.0
26.7
27.0
6.0
9.5
Llama-3.3-70B-Inst
70B
18.4
22.7
52.0
16.0
3.2
GPT-4o-mini [ 16 ]
8B
40.8
35.6
41.0
15.0
7.1
Table 1: Accuracy Results on Search & Agentic Reasoning Benchmarks.
Method
Model Size
Knowledge-Intensive Search
Agentic
Bamboogle
2Wiki
HotpotQA
Musique
GAIA
Qwen2.5-7B-Inst
7B
12.0
23.0
21.0
6.0
3.2
Qwen2.5-14B-Inst
14B
21.6
26.7
20.0
8.0
5.5
Qwen2.5-32B-Inst
32B
24.0
26.7
27.0
6.0
9.5
Llama-3.3-70B-Inst
70B
18.4
22.7
52.0
16.0
3.2
GPT-4o-mini [ 16 ]
8B
40.8
35.6
41.0
15.0
7.1
Table 1: Accuracy Results on Search & Agentic Reasoning Benchmarks.
Method
Model Size
Mathematical Reasoning
Scientific Reasoning
AIME24
AMC23
GameOf24
GPQA
MedQA
Qwen2.5-7B-Inst
7B
6.7
47.5
33.0
34.0
66.0
Qwen2.5-14B-Inst
14B
6.7
60.0
25.0
31.0
75.0
Llama-3.3-70B-Inst
70B
6.7
47.5
31.0
35.0
67.0
Llama-3.1-405B-Inst
405B
26.7
47.5
23.0
30.0
62.0
GPT-4o-mini [ 16 ]
8B
13.3
57.5
16.0
27.0
66.0
Table 2: Accuracy Results on Math & Scientific Reasoning Benchmarks. (7B-Inst: Qwen2.5-7B-Instruct, 7B-Base: Qwen2.5-7B-Base)
Figure 4: Our Case Study Example.
DPO
Multi-Level
Fair
Bamboogle
2Wiki
GAIA
AIME24
✓
61.6%
61.5%
21.3%
30.0%
✓
✓
76.0%
76.0%
36.2%
43.3%
✓
✓
✓
81.6%
81.5%
41.7%
50.0%
Frozen
58.4%
60.0%
17.2%
16.7%
Supervised Fine-Tuning
30.4%
32.7%
6.3%
3.3%
Flow-GRPO
69.6%
77.2%
33.1%
40.0%
Table 3: Effectiveness of Our Proposed Objective.
T
Bamboogle
2Wiki
GAIA
AIME24
3
60.8%
61.0%
21.3%
26.7%
5
69.6%
70.0%
29.9%
36.7%
7
76.8%
77.0%
37.0%
43.3%
10
81.6%
81.5%
41.7%
50.0%
Table 4: Effectiveness of Number of Turns.
γ
Bamboogle
2Wiki
GAIA
AIME24
0.0
76.0%
76.0%
36.2%
43.3%
0.5
76.0%
76.5%
37.0%
46.7%
1.0
78.4%
78.0%
37.8%
46.7%
2.0
81.6%
81.5%
41.7%
50.0%
5.0
75.2%
75.5%
35.4%
43.3%
Table 5: Effectiveness of Hyper-Parameter γ .
Method
Size
Bamboogle
2Wiki
GAIA
AIME24
Frozen
Qwen2.5-3B-Instruct
53.6%
63.0%
14.3%
13.3%
Qwen2.5-7B-Instruct
58.4%
60.0%
17.2%
16.7%
AgentFlow
Qwen2.5-3B-Instruct
68.8%
72.3%
29.1%
20.0%
Qwen2.5-7B-Instruct
69.6%
77.2%
33.1%
40.0%
Ours
Qwen2.5-3B-Instruct
72.0%
72.5%
32.3%
40.0%
Qwen2.5-7B-Instruct
81.6%
81.5%
41.7%
50.0%
Table 6: Effectiveness of Model Size.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
λDPO
λΦ−MPO
Bamboogle
2Wiki
GAIA
AIME24
0.5
1.0
72.8
73.0
33.1
45.0
1.0
1.0
81.6
81.5
41.7
50.0
1.0
0.5
67.2
67.0
26.8
33.3
Appendix
Table 7: Effectiveness of Weighting Parameters (Accuracy).
National Engineering Research Center of Software Engineering, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Key Laboratory of High Confidence Software Technologies, Ministry of Education, Beijing, China +3