Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.
Figures & tables
Figure 1 : 0.06275 0.47843 0.73725V0.12157 0.52157 0.63529e0.18039 0.56471 0.53333r0.23922 0.60784 0.43137i0.29804 0.65098 0.32549F0.35686 0.69412 0.22353i0.41569 0.73725 0.12157n0.47451 0.78039 0.01961e __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: demonstrates continuous self-improvement with the co-evolution of the policy, training curriculum, and judge in an agentic harness framework for embodied reasoning, including autonomous driving and robot navigation tasks. Project Website
Figure 2 : 0.06275 0.47843 0.73725V0.12157 0.52157 0.63529e0.18039 0.56471 0.53333r0.23922 0.60784 0.43137i0.29804 0.65098 0.32549F0.35686 0.69412 0.22353i0.41569 0.73725 0.12157n0.47451 0.78039 0.01961e __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: is an agent harness framework for scaling verification. The Policy Improvement Loop uses a reference-free judge to diagnose recurring failures, adapt the training curriculum, and optimize the policy. When verification limits further progress, the Judge Improvement Loop selectively acquires human guidance and refines the judge through coactive calibration, resolving disagreements in rubric-based judgments. The updated judge guides subsequent data selection and policy optimization, allowing verification capability to evolve with the policy’s failure patterns.
Curriculum Judge
Reward Judge
Reasoning Score (RB) ↑
Reasoning Score (RF) ↑
minADE 6 (m) ↓
ADE (m) ↓
Base Policy Model
None
None
60.56
68.05
1.049
2.139
Reference-Based Judges
Random
LingoJudge
63.45
69.82
1.064
2.133
LingoJudge
LingoJudge
65.75
72.60
1.172
2.242
Random
VeriFine-Judge-RB
64.03
69.95
1.080
2.166
Table 1 : Comparison of Different Judges in the Policy Improvement Loop with a common initialization and matched per-policy training budgets.The curriculum and reward judges refer to the judges for data selection and policy optimization, respectively; reference-free variants use the same for both.
Figure 3 : Judge Performance Comparison. (a) Calibration curves showing the mean judge score within each human-score bin; the gray diagonal denotes ideal calibration. (b) Judge–human alignment measured by Pearson correlation and MAE for teacher and student judges across improvement rounds and reference-based baselines. (c) Test-time scaling is evaluated by the human score of the selected candidate as the candidate budget increases. Shaded regions denote 95% bootstrap CI.
Figure 4 : Self-improvement on Robot Navigation Task. The left part shows the quantitative results in the robot navigation task with significant improvement in final policy performance, and the right part shows the qualitative results with a sequential reasoning example: the updated policy with 0.06275 0.47843 0.73725V0.12157 0.52157 0.63529e0.18039 0.56471 0.53333r0.23922 0.60784 0.43137i0.29804 0.65098 0.32549F0.35686 0.69412 0.22353i0.41569 0.73725 0.12157n0.47451 0.78039 0.01961e __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: __color_backend_reset: corrects the action failures of the old policy.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Curriculum Judge
Reward Judge
Reasoning Score (RB) ↑
Reasoning Score (RF) ↑
minADE 6 (m) ↓
ADE (m) ↓
Random
VeriFine-Judge- R1
58.71
68.90
1.349
2.594
VeriFine-Judge- R1 ( First Round )
67.30
77.51
1.106
2.177
Random
VeriFine-Judge- R2
65.32
76.67
1.119
2.317
VeriFine-Judge- R2 ( Second Round )
70.59
79.61
1.078
2.139
Random
VeriFine-Judge- R3
68.53
81.59
1.053
2.222
VeriFine-Judge- R3 ( Third Round )
71.61
83.13
1.029
2.117
Appendix
Table 2 : Curriculum Ablation across Judge Improvement Rounds. Within each round, random and judge-guided selection share the same reward judge, base-policy initialization, and policy-training budget. Shaded rows use judge-guided selection. All policies are evaluated using the same fixed RB and final RF evaluators.
Figure 5 : Comparison of Backbone Variants of the Teacher Judge Model . All the rubric variants, except the one without coactive calibration, use the same rubric, but with different backbone models. No rubric variant uses a straightforward instruction without detailed rubrics for judgment. (a) Calibration curves showing the mean judge score within each human-score bin; the gray diagonal denotes ideal calibration. (b) Judge–human alignment measured by Pearson correlation and MAE for backbone variants.
Figure 6 : Evolved Judge Performance and Evolved Test Sets (R1 and R2). (a) (c) Calibration curves showing the mean judge score within each human-score bin for two evolved test subsets for the judge test set; the gray diagonal denotes ideal calibration. (b) (d) Judge–human alignment measured by Pearson correlation and MAE for teacher and student judges across improvement rounds and reference-based baselines.
Figure 7 : Teacher–Student Architecture of the Reference-free Rubric Judge. The teacher evaluates policy outputs using visual context, ego state, a driving or robot navigation instruction, and a structured rubric. Its evaluations are distilled into a compact student judge for scalable verification.
With PAI-AV Pretraining
Output Paradigm
Pearson r↑
MAE ↓
Inference Latency (s/sample) ↓
Head
Output
×
Generation
Reasoning + Score
0.72
0.18
2.34
Score Only
0.71
0.19
0.45
Classifier
Overall Score
0.76
0.18
0.07
Rubric Subscores
0.76
0.19
0.07
✓
Generation
Reasoning + Score
0.76
0.16
2.40
Appendix
Table 3 : Comparison of Student Judge Variants. All variants use Qwen3-VL as the backbone. We compare initialization with and without pretraining on our internal dataset, together with generation- and classifier-based output paradigms. Inference latency is averaged over 200 samples at batch size one on NVIDIA H100 GPUs, and the teacher model is accessed through a remote API.
Figure 8 : Scenario Distribution of the Testing Sets in Driving Reasoning . The figure shows the top 21 categories in each test set, which constitute the main body of the sets.
Set
Primary purpose
Construction
Size
P & Ptest
Policy monitoring and testing
Fixed, expert-checked
2134, 2772
H1 & H1test
Initial judge calibration and testing
Broad base-policy rollouts
433, 514
H2 & H2test
Boundary-focused calibration and testing
Updated-policy rollouts
179, 218
J3 & J3test
Final judge validation and testing
H1∪H2 , H1test∪H2test
612, 732
Appendix
Table 4: Summary of the Evaluation and Testing Sets of Driving Reasoning . The query batch Ht contains newly selected cases at judge-improvement iteration t , while Jt denotes the cumulative judge evaluation set.
Set
Evaluation
Test
Policy ( P )
182
220
Initial judge ( H1 )
175
211
R2 added judge cases ( H2 )
37
70
Cumulative judge ( J2 )
212
281
Appendix
Table 5 : Robot Navigation Evaluation and Test Sets.
Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a generator-verifier framework for generalist robot policies for inference-time policy steering and self-improvement. We use a pre-trained generalist robot policy as a generator'' and pair it with a gradient-free visual verifier'' that evaluates actions at inference time. This framework enables inference-time steering that improves policy performance without additional training. We demonstrate that inference-time verification consistently outperforms vanilla generalists without training on additional demonstration data. Additionally, we demonstrate that the verified rollouts provide effective supervision for offline policy improvement: policies fine-tuned on verified self-generated trajectories achieve consistent performance gains. Notably, we find that post-training with verified rollouts achieves comparable efficiency to expert demonstrations, while requiring no human interventions. Our results highlight inference-time verification as a practical and scalable mechanism for improving robotic policies during deployment.
Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.
Diandian Guo, Cong Cao, Fangfang Yuan +3
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
As reinforcement learning continues to scale the training of large language model-based agents, reliably verifying agent behaviors in complex environments has become increasingly challenging. Existing approaches rely on rule-based verifiers or LLM-as-a-Judge models, which struggle to generalize beyond narrow domains. Agent-as-a-Judge addresses this limitation by actively interacting with environments and tools to acquire verifiable evidence, yet its capabilities remain underexplored. We introduce a benchmark AJ-Bench to systematically evaluate Agent-as-a-Judge across three domains-search, data systems, and graphical user interfaces-comprising 155 tasks and 516 annotated trajectories. The benchmark comprehensively assesses judge agents' abilities in information acquisition, state verification, and process verification. Experiments demonstrate consistent performance gains over LLM-as-a-Judge baselines, while also revealing substantial open challenges in agent-based verification. Our data and code are available at https://aj-bench.github.io/.
Wentao Shi, Yu Wang, Yuyang Zhao +8
University of Science and Technology of China · National University of Singapore · Meituan