Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.
Figures & tables
Figure 1 : 0.06275 0.47843 0.73725V0.12157 0.52157 0.63529e0.18039 0.56471 0.53333r0.23922 0.60784 0.43137i0.29804 0.65098 0.32549F0.35686 0.69412 0.22353i0.41569 0.73725 0.12157n0.47451 0.78039 0.01961e __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: demonstrates continuous self-improvement with the co-evolution of the policy, training curriculum, and judge in an agentic harness framework for embodied reasoning, including autonomous driving and robot navigation tasks. Project Website
Figure 2 : 0.06275 0.47843 0.73725V0.12157 0.52157 0.63529e0.18039 0.56471 0.53333r0.23922 0.60784 0.43137i0.29804 0.65098 0.32549F0.35686 0.69412 0.22353i0.41569 0.73725 0.12157n0.47451 0.78039 0.01961e __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: is an agent harness framework for scaling verification. The Policy Improvement Loop uses a reference-free judge to diagnose recurring failures, adapt the training curriculum, and optimize the policy. When verification limits further progress, the Judge Improvement Loop selectively acquires human guidance and refines the judge through coactive calibration, resolving disagreements in rubric-based judgments. The updated judge guides subsequent data selection and policy optimization, allowing verification capability to evolve with the policy’s failure patterns.
Curriculum Judge
Reward Judge
Reasoning Score (RB) ↑
Reasoning Score (RF) ↑
minADE 6 (m) ↓
ADE (m) ↓
Base Policy Model
None
None
60.56
68.05
1.049
2.139
Reference-Based Judges
Random
LingoJudge
63.45
69.82
1.064
2.133
LingoJudge
LingoJudge
65.75
72.60
1.172
2.242
Random
VeriFine-Judge-RB
64.03
69.95
1.080
2.166
Table 1 : Comparison of Different Judges in the Policy Improvement Loop with a common initialization and matched per-policy training budgets.The curriculum and reward judges refer to the judges for data selection and policy optimization, respectively; reference-free variants use the same for both.
Figure 3 : Judge Performance Comparison. (a) Calibration curves showing the mean judge score within each human-score bin; the gray diagonal denotes ideal calibration. (b) Judge–human alignment measured by Pearson correlation and MAE for teacher and student judges across improvement rounds and reference-based baselines. (c) Test-time scaling is evaluated by the human score of the selected candidate as the candidate budget increases. Shaded regions denote 95% bootstrap CI.
Figure 4 : Self-improvement on Robot Navigation Task. The left part shows the quantitative results in the robot navigation task with significant improvement in final policy performance, and the right part shows the qualitative results with a sequential reasoning example: the updated policy with 0.06275 0.47843 0.73725V0.12157 0.52157 0.63529e0.18039 0.56471 0.53333r0.23922 0.60784 0.43137i0.29804 0.65098 0.32549F0.35686 0.69412 0.22353i0.41569 0.73725 0.12157n0.47451 0.78039 0.01961e __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: corrects the action failures of the old policy.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Curriculum Judge
Reward Judge
Reasoning Score (RB) ↑
Reasoning Score (RF) ↑
minADE 6 (m) ↓
ADE (m) ↓
Random
VeriFine-Judge- R1
58.71
68.90
1.349
2.594
VeriFine-Judge- R1 ( First Round )
67.30
77.51
1.106
2.177
Random
VeriFine-Judge- R2
65.32
76.67
1.119
2.317
VeriFine-Judge- R2 ( Second Round )
70.59
79.61
1.078
2.139
Random
VeriFine-Judge- R3
68.53
81.59
1.053
2.222
VeriFine-Judge- R3 ( Third Round )
71.61
83.13
1.029
2.117
Appendix
Table 2 : Curriculum Ablation across Judge Improvement Rounds. Within each round, random and judge-guided selection share the same reward judge, base-policy initialization, and policy-training budget. Shaded rows use judge-guided selection. All policies are evaluated using the same fixed RB and final RF evaluators.
Figure 5 : Comparison of Backbone Variants of the Teacher Judge Model . All the rubric variants, except the one without coactive calibration, use the same rubric, but with different backbone models. No rubric variant uses a straightforward instruction without detailed rubrics for judgment. (a) Calibration curves showing the mean judge score within each human-score bin; the gray diagonal denotes ideal calibration. (b) Judge–human alignment measured by Pearson correlation and MAE for backbone variants.
Figure 6 : Evolved Judge Performance and Evolved Test Sets (R1 and R2). (a) (c) Calibration curves showing the mean judge score within each human-score bin for two evolved test subsets for the judge test set; the gray diagonal denotes ideal calibration. (b) (d) Judge–human alignment measured by Pearson correlation and MAE for teacher and student judges across improvement rounds and reference-based baselines.
Figure 7 : Teacher–Student Architecture of the Reference-free Rubric Judge. The teacher evaluates policy outputs using visual context, ego state, a driving or robot navigation instruction, and a structured rubric. Its evaluations are distilled into a compact student judge for scalable verification.
With PAI-AV Pretraining
Output Paradigm
Pearson r↑
MAE ↓
Inference Latency (s/sample) ↓
Head
Output
×
Generation
Reasoning + Score
0.72
0.18
2.34
Score Only
0.71
0.19
0.45
Classifier
Overall Score
0.76
0.18
0.07
Rubric Subscores
0.76
0.19
0.07
✓
Generation
Reasoning + Score
0.76
0.16
2.40
Appendix
Table 3 : Comparison of Student Judge Variants. All variants use Qwen3-VL as the backbone. We compare initialization with and without pretraining on our internal dataset, together with generation- and classifier-based output paradigms. Inference latency is averaged over 200 samples at batch size one on NVIDIA H100 GPUs, and the teacher model is accessed through a remote API.
Figure 8 : Scenario Distribution of the Testing Sets in Driving Reasoning . The figure shows the top 21 categories in each test set, which constitute the main body of the sets.
Set
Primary purpose
Construction
Size
P & Ptest
Policy monitoring and testing
Fixed, expert-checked
2134, 2772
H1 & H1test
Initial judge calibration and testing
Broad base-policy rollouts
433, 514
H2 & H2test
Boundary-focused calibration and testing
Updated-policy rollouts
179, 218
J3 & J3test
Final judge validation and testing
H1∪H2 , H1test∪H2test
612, 732
Appendix
Table 4: Summary of the Evaluation and Testing Sets of Driving Reasoning . The query batch Ht contains newly selected cases at judge-improvement iteration t , while Jt denotes the cumulative judge evaluation set.
Set
Evaluation
Test
Policy ( P )
182
220
Initial judge ( H1 )
175
211
R2 added judge cases ( H2 )
37
70
Cumulative judge ( J2 )
212
281
Appendix
Table 5 : Robot Navigation Evaluation and Test Sets.
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China