VERA: Scaling Verifiable Environments for Agentic co-Evolution
Authors: Junqi Liu, Yongyang Pan, Zhuosong Jiang, Dongbai Li, Bo Zhang, Xitong Ling, Sheng Wang, Hanrong Ye, +9 more
Organizations: University of California, Santa Cruz · University of Illinois Urbana-Champaign · National University of Singapore · Tsinghua University · Xian Jiaotong University · University of Pennsylvania · Nvidia
Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. To address the challenges in stable training, we present VERA, which builds such environments at scale and lets agents evolve on them. VERA builds these environments from initial trajectories: an agent writes rubrics, executable checks, a judge verifies each sandbox, and only those that pass enter the training bank. On these environments, VERA alternates between two updates: train the model with rubric rewards, or edit the harness skills. We also create a verifier which gates model checkpoints and harness edits using explicit development-set acceptance criteria. This attribution distinguishes VERA's co-evolution from single-axis baselines: its updates target not only the cause but the outcome. With an open-source corpus of 9,000+ long-horizon verifiable environments, a 9B model paired with its co-evolved agent beats the strongest baseline by 10.3 and 13.0 points in the two domains. At 27B, it surpasses the baseline on AutoCoWorkBench (71.6) and AutoMedBench (80.7), transfers to unseen workflows, and retains general capabilities.
Figures & tables
Figure 1: VERA enables model–harness co-evolution through verifiable environments. Top: Auto-generated rubrics, Agent Judge workspace checks, and container replay verify environments before scaling. Bottom left: Attribution uses stage-wise scores and traces to produce an LLM Report selecting training stages and a Harness Report proposing skill additions, merges, or removals. The Reward Verifier checks rubric rewards; the Harness Verifier tests skill edits. Updates alternate with the other component frozen, carrying accepted changes forward. Right: CoWork and MedResearch scores across rounds. Dashed lines denote SFT baselines; labels identify target stages and skill edits. Fallback restores a previously accepted checkpoint.
Figure 2: VERA improves agents across model sizes and achieves competitive performance with a 27B backbone. (a-b) Co-evolution improves on the base agents at 4B, 9B, and 27B; dotted lines mark the best frontier reference scores. (c-d) VERA-27B achieves the highest Overall scores among the compared agents on AutoCoWorkBench and AutoMedBench. (e-f) On Automation-Bench (public 600) ( Zapier, 2026 ) and AgentClinic ( Schmidgall and others, 2024 ) , VERA-27B improves over its base agent by 19.2 and 32.5 percentage points, respectively, ranking second and third among the compared systems. These comparisons evaluate complete agents with their configured productivity harnesses, rather than model weights alone. Green denotes VERA; gray denotes comparison models. Harness and reasoning-effort settings appear in Appendix B.1 (Table 9 ).
Figure 3: Auto-Rubric makes long workflows restartable and verifiable. (1) Training sandbox construction. Each sandbox combines a verified environment (tools, workspace, and initial state), a task query, and the interaction history needed to resume the task. (2) Rubric generation and evidence verification. A coding agent generates binary rubric items from skill configurations, tool API documentation, and execution logs. Each item defines a success criterion scored in {0,1} from observable evidence. Executable checks validate skill and tool schemas. The Agent Judge inspects the trajectory and workspace to assess argument validity, actual execution effects, artifact evidence, and output quality in the task context. Failed executable checks override judge scores. (3) Stage-wise and end-to-end environment synthesis. Benchmark trajectories are segmented into five stages: plan, setup, verify, execute, and submit (S1–S5). Stage-wise environments restore each stage’s starting state and attach the corresponding query, history, and rubrics, allowing that stage to be practiced and scored independently. End-to-end environments preserve the full workflow while an LLM surrogate simulates responses from time-consuming tool calls ( Ren et al., 2025 ) . Both types must pass executable checks and Agent Judge verification before entering the training bank. Numbers beside the environments illustrate rubric scores.
Evolution
Method
Weights
Harness
Reward
Verification
Harness
ADAS ( Hu et al., 2025 )
✗
✓
✗
✗
Ours (VERA Harness)
✗
✓
✗
✓
LLM
SFT ( Ouyang et al., 2022 )
✓
✗
✗
✗
GRPO ( Shao et al., 2024 )
✓
✗
✓
✓
RaR ( Gunjal et al., 2025 )
✓
✗
✓
✓
OPD ( Agarwal et al., 2024 )
✓
✗
✗
✗
Table 1: VERA jointly evolves the LLM and harness with rubric-based rewards and verification. Weights and Harness indicate updated components. Reward and Verification indicate rubric-based training rewards and verified-rubric checks in our implementations.
Method
AutoCoWork Bench
SWE-Bench Verified
AutoMed Bench
MedXpertQA Text
O
A
T
P@1
O
A
T
P@1
Base LLM + benchmark default harness
Qwen3.5-9B ( Qwen Team, 2026b )
4.2
0.0
8.3
41.8
19.9
11.3
28.4
35.4
Qwen3.5-9B ( Qwen Team, 2026b ) + Prompt
12.0
8.0
16.0
44.0
22.1
35.3
8.9
40.8
Harness Evolving
ADAS ( Hu et al., 2025 )
20.7
22.3
19.1
52.4
55.5
61.6
49.5
42.0
Table 2: VERA achieves the best Overall and Pass@1 scores across both domains. All methods use Qwen3.5-9B ( Qwen Team, 2026b ) , with separate CoWork and MedResearch agents. O/A/T denote Overall/Agentic/Task, with Overall averaging Agentic and Task. SWE-Bench Verified ( Jimenez et al., 2024 ) and MedXpertQA Text ( Zuo et al., 2025 ) report Pass@1. Bold and underlined values indicate the best and second-best scores per column, respectively. Configuration details appear in Appendix F.5 .
AutoCoWorkBench
AutoMedBench
Agent
O
A
T
Agent
O
A
T
Base LLM + benchmark default harness
Qwen3.8-27B ( Qwen Team, 2026c )
63.8
60.8
66.7
Qwen3.8-27B ( Qwen Team, 2026c )
65.1
81.6
48.6
GPT-5.6-Sol ( OpenAI, 2026 )
72.6
76.8
68.3
GPT-5.6-Sol ( OpenAI, 2026 )
75.1
80.1
70.0
Claude-Opus-4.8 ( Anthropic, 2026 )
66.6
70.7
62.5
Claude-Opus-4.8 ( Anthropic, 2026 )
81.9
88.1
75.8
Base LLM + VERA Harness
Table 3: VERA-27B comes within 1.2 Overall points of the best frontier baseline in each domain. O/A/T denote Overall/Agentic/Task. Frontier baselines use benchmark-default runners, while VERA variants use the model and harness configurations specified by the row groups. This table complements Figure 2 , which evaluates agents with productivity harnesses; rankings are specific to the evaluated configurations. Bold and underlined values indicate the best and second-best scores per column, respectively. Agent and harness configurations are detailed in Appendix B .
Model
AIME 2026
ALFWorld
GPQA-Diamond
IF-Bench
Qwen3.5-4B ( Qwen Team, 2026a )
89.7
32.1
76.2
59.2
VERA-CoWork-4B
33.3 ( − 56.4)
38.1 (+6.0)
64.3 ( − 11.9)
65.5 (+6.3)
VERA-Med-4B
30.0 ( − 59.7)
18.2 ( − 13.9)
61.1 ( − 15.1)
34.5 ( − 24.7)
Qwen3.5-9B ( Qwen Team, 2026b )
92.5
74.6
81.7
66.7
VERA-CoWork-9B
36.7 ( − 55.8)
53.2 ( − 21.4)
57.1 ( − 24.6)
59.8 ( − 6.9)
VERA-Med-9B
76.7 ( − 15.8)
68.4 ( − 6.2)
81.2 ( − 0.5)
52.0 ( − 14.7)
Table 4: General-capability retention. Parentheses show score changes from the corresponding backbone (green: gains; red: drops).
Figure 4: Environment preparation dominates the estimated RSI cost in Medical Research. Stacked areas show environment scaling and curation, training infrastructure, and verification and attribution costs. Dashed lines mark the start of model-harness co-evolution at 75% of environment preparation and the completion of environment preparation. The normalized timeline and smooth curves illustrate the schedule rather than measured spending over time.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Collection
Size
Role and boundary
Lite evaluation core
40 tasks
Complete episodes; the official Lite evaluation denominator.
Training stage pool
56 tasks
Stage-specific practice; excluded from evaluation.
Gate-1000
Separate pool
Training or gated development; restricted evaluators stay outside agent workspaces.
PawBench-derived pack
46 tasks
Separate CoWork pack with answer-free workspaces and offline evaluators; excluded from the Lite core.
Appendix
Table 5: AutoCoWorkBench separates evaluation tasks from training and gating pools. Counts refer to the dataset repository’s task manifests.
Track
Cases
Task metric
Classification
100
Accuracy
Synthesis
20
Mean SSIM
Detection
100
mAP at IoU 0.5
Segmentation
40
Macro mean Dice
Visual question answering
2,005
Accuracy
Report generation
100
Lightweight report quality
Appendix
Table 6: AutoMedBench Lite groups cases into seven medical workflows. The Task metric evaluates the output artifact for each track.
Benchmark
Domain
Runner / evaluation protocol
AutoCoWorkBench
CoWork
mini-SWE-agent
SWE-Bench Verified
CoWork
Official SWE-bench evaluation harness
Automation-Bench
CoWork
Official Zapier runner
AutoMedBench
Med
react loop with model API calls
MedXpertQA Text
Med
Direct prompting with the official scorer
AgentClinic
Med
Official clinic simulator
Appendix
Table 7: Benchmark domains and evaluation protocols.
Setting
Configuration
Domains
CoWork and MedResearch
Agent harness
VERA harness with domain-specific skills; Codex v0.153.4
Thinking
Enabled and preserved; reasoning effort xhigh
Context window
262,144 tokens
Compaction trigger
200,000 context tokens
Compaction reserve
16,384 tokens
Appendix
Table 8: Shared VERA harness configuration across CoWork and MedResearch evaluations. Benchmark-specific limits take precedence where required.
Candidate
Harness
Serving
Effort
VERA-CoWork-27B
VERA-Codex v0.153.4
SGLang
xhigh
VERA-Med-27B
VERA-Codex v0.153.4
SGLang
xhigh
Qwen3.5-4B ( Qwen Team, 2026a )
Codex v0.153.4
API
xhigh
Qwen3.5-9B ( Qwen Team, 2026b )
Codex v0.153.4
API
xhigh
Qwen3.8-27B ( Qwen Team, 2026c )
Codex v0.153.4
API
xhigh
DeepSeek-V4-Pro ( DeepSeek-AI, 2026 )
Codex v0.153.4
API
xhigh
Appendix
Table 9: Evaluation candidate configurations. Model sources: Qwen ( Qwen Team, 2026a ; Qwen Team, 2026b ; Qwen Team, 2026c ) , DeepSeek ( DeepSeek-AI, 2026 ) , GPT ( OpenAI, 2026 ) , GLM ( Z.ai, 2026 ) , and Claude ( Anthropic, 2026 ) . VERA serving uses SGLang ( Zheng et al., 2023b ) . Harness versions are fixed throughout evaluation. "xhigh" is the highest thinking effort setting, due to API providers’ configuration.
Stage
Shared checks
Domain-specific examples
S1: Plan
Record constraints, required inputs, an ordered method, and observable success and validation conditions.
Bound command and filesystem scope; identify the defect and repair tests; specify data splits and the target metric.
S2: Setup
Inspect the workspace, confirm inputs and runtime capabilities, and record a baseline without unauthorized changes.
Check command dependencies; capture repository state and failing-test evidence; verify dataset schemas and execution configuration.
S3: Verify
Run a bounded representative trial, bind observations to tool records, and check readiness and recovery.
Inspect pilot exit codes; execute a focused reproduction or integration test; recompute pilot metrics and check for data leakage.
S4: Execute
Produce the required outputs using the validated method, resolve errors, and preserve unrelated state.
Check command outputs; apply scoped patches and run regression tests; complete the learning pipeline and validate prediction artifacts.
S5: Submit
Reopen deliverables, run final checks, and support completion claims with artifact and execution evidence.
Validate output manifests; match the submitted patch to the final diff; reload model artifacts and verify submission schemas.
Appendix
Table 10: CoWork rubrics combine shared workflow checks with domain-specific evidence. Each stage has six shared and four domain-specific binary items. Entries below summarize representative checks rather than list all ten items.
Stage
Representative criteria
Observable evidence
S1: Plan
Define the research objective, method, inputs, deliverable, uncertainty, and stopping condition before execution.
A task-bound plan recorded before setup, with an explicit output contract and bounded procedure.
S2: Setup
Acquire permitted inputs, prepare the runtime, inspect relevant evidence, and preserve provenance.
Asset and dependency records, successful loading or inspection, and task-relevant observations such as image geometry.
S3: Verify
Run and inspect a bounded trial before full execution; validate its output and address observed failures.
Pilot tool records and artifacts; checks of labels, coordinates, mask geometry, or the task’s analysis format.
S4: Execute
Complete the declared workload, save the required outputs, and retain their links to the inputs, method, and limitations.
Materialized artifacts, item-count checks, schema validation, and provenance linking full outputs to the inspected inputs and pilot.
S5: Submit
Validate the final artifact, submit one consistent result, and report only supported status and uncertainty.
Independent artifact checks, an accepted submission record, and a final response consistent with that record.
Appendix
Table 11: MedResearch rubrics bind stage progress to observable medical-workflow evidence. Base stage tables contain six equally weighted binary items. Examples summarize medical-image and research workflows; exact checks depend on the task family.
Setting
CoWork
MedResearch
Policy
Qwen3.8-27B ( Qwen Team, 2026c )
Training objective
Full-parameter GRPO
Precision
BF16
Optimizer
Adam
Batch
12 prompts × 8 rollouts
8 prompts × 8 rollouts
Context limit
49,152 tokens
262,144 tokens
Appendix
Table 12: VERA training settings and compute cost. Both domains use full-parameter on-policy GRPO in BF16. Compute costs are approximate GPU-hours.
Table 13: Baseline training and search settings. Fixed-harness RL baselines use 500 updates, matching five 100-update VERA rounds. The COS-PLAY adaptation uses 200 decision-policy updates after SFT-200 initialization. Details appear in Appendix F .
Method
O
A
T
Base + VERA LLM + VERA Harness (Ours)
31.0 ★
45.3 ★
16.7
Base + ADAS harness ( Hu et al., 2025 )
20.7 ★
22.3
19.1
Base + COS-PLAY skill bank ( Wu et al., 2026b )
19.9
18.2
21.6 ★
Base + OPD ( Agarwal et al., 2024 )
19.8
15.6
24.0 ★
Base + VERA LLM (harness frozen)
18.9
23.6 ★
14.2
Base + VERA Harness (LLM frozen)
18.2
21.3
15.1
Appendix
Table 14: CoWork line, sorted by Overall. All methods use Qwen3.5-9B ( Qwen Team, 2026b ) . Base and Base Prompt use the benchmark-default runner; other configurations follow Appendix F.1 . O/A/T denote Overall/Agentic/Task on AutoCoWorkBench.
Method
O
A
T
Base + VERA LLM + VERA Harness (Ours)
69.1 ★
76.8 ★
61.3 ★
Base + VERA Harness (LLM frozen)
56.1 ★
63.4 ★
48.8
Base + ADAS harness ( Hu et al., 2025 )
55.5
61.6
49.5 ★
Base + VERA LLM (harness frozen)
43.3
50.6
36.1
Base + OPD ( Agarwal et al., 2024 )
41.2
51.8
30.6
Base + GRPO ( Shao et al., 2024 )
36.3
47.0
25.7
Appendix
Table 15: MedResearch line, sorted by Overall. All methods use Qwen3.5-9B ( Qwen Team, 2026b ) . Base and Base Prompt use the benchmark-default runner; other configurations follow Appendix F.1 . O/A/T denote Overall/Agentic/Task on AutoMedBench.
Figure 5: The LLM Report directs model training; the Harness Report directs skill edits. This illustrative CoWork example follows the update and acceptance rules in Section 2.1 ; report excerpts and task traces are synthetic. The worked-example presentation follows Figure 2 of COS-PLAY ( Wu et al., 2026b ) .
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Caiqi Zhang, Rujun Han, Zifeng Wang +4
University of Cambridge · Google Cloud AI Research
Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model Context Protocol (MCP) and broader agent skills offer a unified interface for connecting agents with scalable real-world services, but training robust agents remains limited by the lack of realistic environments and principled mechanisms for life-long learning. In this paper, we present \textbf{Agent-World}, a self-evolving training arena for advancing general agent intelligence through scalable environments. Agent-World has two main components: (1) Agentic Environment-Task Discovery, which autonomously explores topic-aligned databases and executable tool ecosystems from thousands of real-world environment themes and synthesizes verifiable tasks with controllable difficulty; and (2) Continuous Self-Evolving Agent Training, which combines multi-environment reinforcement learning with a self-evolving agent arena that automatically identifies capability gaps through dynamic task synthesis and drives targeted learning, enabling the co-evolution of agent policies and environments. Across 23 challenging agent benchmarks, Agent-World-8B and 14B consistently outperforms strong proprietary models and environment scaling baselines. Further analyses reveal scaling trends in relation to environment diversity and self-evolution rounds, offering insights for building general agent intelligence.
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.