Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while single-pair Actor-Critic collaboration refines only one candidate. We introduce SEPAL, which assigns three private Actor-Critic teams to direct reasoning, evidence grounding, and verification. Role-specific training gives the teams different reasoning objectives beyond sampling variation. Each Critic guides revisions within its own team, preventing feedback from carrying errors across candidates. Once revision ends, majority voting combines only the final answers, keeping the reasoning histories separate until the decision. Across five open-weight backbones and five question-answering benchmarks, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor-Critic pair, with improvements across all five backbones. Code is available at https://github.com/zhansan114514/SEPAL.
Figures & tables
Figure 1: SEPAL in one view. A question enters three isolated role pairs, each refined through four private Critic-guided revisions (R1–R4). Only final parsed answers cross team boundaries. A fixed majority vote returns the prediction and uses Direct for split decisions.
Figure 2: Training the private teams. Role SFT initializes Actors; Critics start from the base model. Feedback is valued by Actor continuation correctness. Critic DPO precedes Actor preference construction and Actor DPO. Each role follows this sequence separately before inference in Figure 1 .
Dataset
Split
N
Use
MMLU
test
14,042
in-domain
BoolQ
validation
3,270
transfer
BBH
22-category stratified
1,260
transfer
SciQ
test
1,000
transfer
ARC
Easy + Challenge test
3,548
transfer
Table 1: Evaluation datasets. Transfer rows supply no training examples.
Model
Dataset
Direct
Debate
SoM-2
SoM-4
ACC
SEPAL
Llama-3-8B
BoolQ
77.34
76.76
78.98
78.74
76.54
76.70
MMLU
62.41
63.55
63.39
63.33
65.00
66.93
BBH
49.52
50.00
50.52
51.31
53.17
56.67
SciQ
91.90
92.00
92.45
92.03
91.50
93.40
ARC
88.30
88.92
89.04
88.65
89.04
90.78
Macro
73.89
74.25
74.87
74.81
75.05
76.90
Table 2: Accuracy (%). ACC is the matched single-team implementation and SEPAL uses three teams. Black bold marks each row’s best value; pale blue identifies SEPAL throughout. Macro averages the five datasets.
Figure 3: Decision diagnostics across all 25 cells. (a) Accuracy change from ACC to SEPAL. (b) Mean role, vote, and oracle-any-role accuracy. (c) Majority coverage and unanimity by dataset, averaged over backbones.
Model
SFT
Base-C
Trained-C
Full-R0
No-SFT
Full-R4
Llama-3-8B
71.95
76.35
76.72
72.27
77.79
76.90
Qwen2.5-3B
75.57
76.71
77.79
75.98
75.98
77.21
Gemma-2-2B
69.58
71.96
72.31
69.70
71.98
72.52
Phi-4-mini
74.37
80.29
81.03
75.41
80.41
80.90
Mistral-7B
70.20
73.38
73.66
69.60
73.65
74.03
Mean
72.34
75.74
76.30
72.59
75.96
76.31
Table 3: Macro accuracy (%) over five datasets. Black bold marks the best variant for each backbone; pale blue identifies Full-R4.
Figure 4: Component ablations and revision rounds. Left: observed macro-accuracy contrasts averaged over five backbones. Right: majority-vote macro accuracy by Actor round, with the first Critic-conditioned revision at R1.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Checkpoint identifier
Llama-3-8B
meta-llama/Meta-Llama-3-8B-Instruct
Qwen2.5-3B
Qwen/Qwen2.5-3B-Instruct
Gemma-2-2B
google/gemma-2-2b-it
Phi-4-mini
microsoft/Phi-4-mini-instruct
Mistral-7B
mistralai/Mistral-7B-Instruct-v0.3
Appendix
Table 4: Public checkpoint identities used in every reported run.
Setting
Role SFT
Preference / DPO
Evaluation
Source split
MMLU auxiliary train
MMLU validation
benchmark-specific
Source questions
10,000
1,531
all configured
Generation temperature
{0.4,0.7,1.0}
0.7
0.7
Top- p
0.9
0.9
0.9
Maximum new tokens
1,024
1,024
1,024
Epochs
1
3
Not used
Appendix
Table 5: Resolved settings for role initialization, preference optimization, and evaluation.
Model
SFT / role
Critic pairs
Actor pairs
Llama-3-8B
7,847
827–2,041
437–828
Qwen2.5-3B
8,313
429–771
960–1,032
Gemma-2-2B
6,902
494–593
704–927
Phi-4-mini
7,755
276–533
880–975
Mistral-7B
7,005
1,077–2,223
2,811–4,813
Appendix
Table 6: Observed optimization-data yields. Preference entries give the role-wise minimum–maximum under the fixed margin rule.
Model
Dataset
N
ACC
SEPAL
Δ
Llama-3-8B
BoolQ
3,270
76.54
76.70
+0.15
MMLU
14,042
65.00
66.93
+1.93
BBH
1,260
53.17
56.67
+3.49
SciQ
1,000
91.50
93.40
+1.90
ARC
3,548
89.04
90.78
+1.75
Qwen2.5-3B
BoolQ
3,270
73.30
77.71
+4.40
Appendix
Table 7: Complete matched comparison (accuracy, %). The five rows within a model use different evaluation sets but the same trained policy family and decision rule.
Model
Variant
BoolQ
MMLU
BBH
SciQ
ARC
Macro
Llama-3-8B
SFT-only
73.82
60.58
49.84
88.20
87.29
71.95
SFT+Base-C
78.47
64.48
56.03
92.90
89.85
76.35
SFT+Trained-C
78.99
65.57
55.71
92.80
90.53
76.72
Full-R0
71.47
62.15
51.35
88.80
87.57
72.27
No-SFT
79.94
67.05
57.62
93.80
90.53
77.79
Full-R4
76.70
66.93
56.67
93.40
90.78
76.90
Appendix
Table 8: Complete component matrix (accuracy, %). Black bold marks the best value per dataset and backbone; Macro weights all five datasets equally.
Model
R0
R1
R2
R3
R4
Llama-3-8B
72.27
76.12
76.54
77.04
76.90
Qwen2.5-3B
75.98
76.97
77.25
77.35
77.21
Gemma-2-2B
69.70
72.43
72.65
72.32
72.52
Phi-4-mini
75.41
80.73
80.93
80.78
80.90
Mistral-7B
69.60
73.28
73.79
73.89
74.03
Mean
72.59
75.91
76.23
76.27
76.31
Appendix
Table 9: Round-wise macro accuracy (%). R1 is the first Critic-conditioned revision. Pale blue marks the fixed endpoint; black bold marks the best round.
Model
Dataset
Direct
Evidence
Verification
SEPAL
Oracle-any
Llama-3-8B
BoolQ
74.19
71.56
80.52
76.70
88.26
MMLU
64.86
64.56
65.23
66.93
80.98
BBH
53.81
52.06
56.19
56.67
75.63
SciQ
92.80
91.80
91.20
93.40
96.60
ARC
90.02
89.04
89.29
90.78
95.29
Qwen2.5-3B
BoolQ
68.96
79.51
77.49
77.71
88.13
Appendix
Table 10: Per-role, voted, and oracle-any-role accuracy (%). Pale blue marks the actual SEPAL decision; mint marks diagnostic oracle headroom.
Role pair
Agreement coverage
Accuracy when agreeing
Direct + Evidence
79.17
81.80
Direct + Verification
78.70
82.04
Evidence + Verification
79.22
82.09
Appendix
Table 11: Pairwise final-answer agreement averaged over all 25 cells (%). Conditional accuracy evaluates the shared answer only on agreeing examples.
Model
Dataset
Majority
Unanimous
Oracle-any
Fallback
Parsed
Llama-3-8B
BoolQ
99.42
73.06
88.26
0.58
99.79
MMLU
93.51
58.48
80.98
6.49
99.94
BBH
89.52
43.97
75.63
10.48
99.76
SciQ
99.20
89.10
96.60
0.80
100.00
ARC
98.70
85.96
95.29
1.30
100.00
Qwen2.5-3B
BoolQ
99.88
71.90
88.13
0.12
100.00
Appendix
Table 12: Complete final-round decision diagnostics (%). Majority is the fraction with a valid two-of-three answer; Unanimous requires all three normalized answers to agree; Fallback is the complement of Majority.