VERA: Scaling Verifiable Environments for Agentic co-Evolution
Organizations: University of California, Santa Cruz · University of Illinois Urbana-Champaign · National University of Singapore · Tsinghua University · Xian Jiaotong University · University of Pennsylvania · Nvidia
Abstract
Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. To address the challenges in stable training, we present VERA, which builds such environments at scale and lets agents evolve on them. VERA builds these environments from initial trajectories: an agent writes rubrics, executable checks, a judge verifies each sandbox, and only those that pass enter the training bank. On these environments, VERA alternates between two updates: train the model with rubric rewards, or edit the harness skills. We also create a verifier which gates model checkpoints and harness edits using explicit development-set acceptance criteria. This attribution distinguishes VERA's co-evolution from single-axis baselines: its updates target not only the cause but the outcome. With an open-source corpus of 9,000+ long-horizon verifiable environments, a 9B model paired with its co-evolved agent beats the strongest baseline by 10.3 and 13.0 points in the two domains. At 27B, it surpasses the baseline on AutoCoWorkBench (71.6) and AutoMedBench (80.7), transfers to unseen workflows, and retains general capabilities.
Figures & tables
| Evolution | Method | Weights | Harness | Reward | Verification |
| Harness | ADAS ( Hu et al., 2025 ) | ✗ | ✓ | ✗ | ✗ |
| Ours (VERA Harness) | ✗ | ✓ | ✗ | ✓ | |
| LLM | SFT ( Ouyang et al., 2022 ) | ✓ | ✗ | ✗ | ✗ |
| GRPO ( Shao et al., 2024 ) | ✓ | ✗ | ✓ | ✓ | |
| RaR ( Gunjal et al., 2025 ) | ✓ | ✗ | ✓ | ✓ | |
| OPD ( Agarwal et al., 2024 ) | ✓ | ✗ | ✗ | ✗ |
| Method | AutoCoWork Bench | SWE-Bench Verified | AutoMed Bench | MedXpertQA Text | ||||
| O | A | T | P@1 | O | A | T | P@1 | |
| Base LLM + benchmark default harness | ||||||||
| Qwen3.5-9B ( Qwen Team, 2026b ) | 4.2 | 0.0 | 8.3 | 41.8 | 19.9 | 11.3 | 28.4 | 35.4 |
| Qwen3.5-9B ( Qwen Team, 2026b ) + Prompt | 12.0 | 8.0 | 16.0 | 44.0 | 22.1 | 35.3 | 8.9 | 40.8 |
| Harness Evolving | ||||||||
| ADAS ( Hu et al., 2025 ) | 20.7 | 22.3 | 19.1 | 52.4 | 55.5 | 61.6 | 49.5 | 42.0 |
| AutoCoWorkBench | AutoMedBench | ||||||
|---|---|---|---|---|---|---|---|
| Agent | O | A | T | Agent | O | A | T |
| Base LLM + benchmark default harness | |||||||
| Qwen3.8-27B ( Qwen Team, 2026c ) | 63.8 | 60.8 | 66.7 | Qwen3.8-27B ( Qwen Team, 2026c ) | 65.1 | 81.6 | 48.6 |
| GPT-5.6-Sol ( OpenAI, 2026 ) | 72.6 | 76.8 | 68.3 | GPT-5.6-Sol ( OpenAI, 2026 ) | 75.1 | 80.1 | 70.0 |
| Claude-Opus-4.8 ( Anthropic, 2026 ) | 66.6 | 70.7 | 62.5 | Claude-Opus-4.8 ( Anthropic, 2026 ) | 81.9 | 88.1 | 75.8 |
| Base LLM + VERA Harness | |||||||
| Model | AIME 2026 | ALFWorld | GPQA-Diamond | IF-Bench |
|---|---|---|---|---|
| Qwen3.5-4B ( Qwen Team, 2026a ) | 89.7 | 32.1 | 76.2 | 59.2 |
| VERA-CoWork-4B | 33.3 ( 56.4) | 38.1 (+6.0) | 64.3 ( 11.9) | 65.5 (+6.3) |
| VERA-Med-4B | 30.0 ( 59.7) | 18.2 ( 13.9) | 61.1 ( 15.1) | 34.5 ( 24.7) |
| Qwen3.5-9B ( Qwen Team, 2026b ) | 92.5 | 74.6 | 81.7 | 66.7 |
| VERA-CoWork-9B | 36.7 ( 55.8) | 53.2 ( 21.4) | 57.1 ( 24.6) | 59.8 ( 6.9) |
| VERA-Med-9B | 76.7 ( 15.8) | 68.4 ( 6.2) | 81.2 ( 0.5) | 52.0 ( 14.7) |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Collection | Size | Role and boundary |
|---|---|---|
| Lite evaluation core | 40 tasks | Complete episodes; the official Lite evaluation denominator. |
| Training stage pool | 56 tasks | Stage-specific practice; excluded from evaluation. |
| Gate-1000 | Separate pool | Training or gated development; restricted evaluators stay outside agent workspaces. |
| PawBench-derived pack | 46 tasks | Separate CoWork pack with answer-free workspaces and offline evaluators; excluded from the Lite core. |
| Track | Cases | Task metric |
|---|---|---|
| Classification | 100 | Accuracy |
| Synthesis | 20 | Mean SSIM |
| Detection | 100 | mAP at IoU 0.5 |
| Segmentation | 40 | Macro mean Dice |
| Visual question answering | 2,005 | Accuracy |
| Report generation | 100 | Lightweight report quality |
| Benchmark | Domain | Runner / evaluation protocol |
| AutoCoWorkBench | CoWork | mini-SWE-agent |
| SWE-Bench Verified | CoWork | Official SWE-bench evaluation harness |
| Automation-Bench | CoWork | Official Zapier runner |
| AutoMedBench | Med | react loop with model API calls |
| MedXpertQA Text | Med | Direct prompting with the official scorer |
| AgentClinic | Med | Official clinic simulator |
| Setting | Configuration |
|---|---|
| Domains | CoWork and MedResearch |
| Agent harness | VERA harness with domain-specific skills; Codex v0.153.4 |
| Thinking | Enabled and preserved; reasoning effort xhigh |
| Context window | 262,144 tokens |
| Compaction trigger | 200,000 context tokens |
| Compaction reserve | 16,384 tokens |
| Candidate | Harness | Serving | Effort |
|---|---|---|---|
| VERA-CoWork-27B | VERA-Codex v0.153.4 | SGLang | xhigh |
| VERA-Med-27B | VERA-Codex v0.153.4 | SGLang | xhigh |
| Qwen3.5-4B ( Qwen Team, 2026a ) | Codex v0.153.4 | API | xhigh |
| Qwen3.5-9B ( Qwen Team, 2026b ) | Codex v0.153.4 | API | xhigh |
| Qwen3.8-27B ( Qwen Team, 2026c ) | Codex v0.153.4 | API | xhigh |
| DeepSeek-V4-Pro ( DeepSeek-AI, 2026 ) | Codex v0.153.4 | API | xhigh |
| Stage | Shared checks | Domain-specific examples |
|---|---|---|
| S1: Plan | Record constraints, required inputs, an ordered method, and observable success and validation conditions. | Bound command and filesystem scope; identify the defect and repair tests; specify data splits and the target metric. |
| S2: Setup | Inspect the workspace, confirm inputs and runtime capabilities, and record a baseline without unauthorized changes. | Check command dependencies; capture repository state and failing-test evidence; verify dataset schemas and execution configuration. |
| S3: Verify | Run a bounded representative trial, bind observations to tool records, and check readiness and recovery. | Inspect pilot exit codes; execute a focused reproduction or integration test; recompute pilot metrics and check for data leakage. |
| S4: Execute | Produce the required outputs using the validated method, resolve errors, and preserve unrelated state. | Check command outputs; apply scoped patches and run regression tests; complete the learning pipeline and validate prediction artifacts. |
| S5: Submit | Reopen deliverables, run final checks, and support completion claims with artifact and execution evidence. | Validate output manifests; match the submitted patch to the final diff; reload model artifacts and verify submission schemas. |
| Stage | Representative criteria | Observable evidence |
|---|---|---|
| S1: Plan | Define the research objective, method, inputs, deliverable, uncertainty, and stopping condition before execution. | A task-bound plan recorded before setup, with an explicit output contract and bounded procedure. |
| S2: Setup | Acquire permitted inputs, prepare the runtime, inspect relevant evidence, and preserve provenance. | Asset and dependency records, successful loading or inspection, and task-relevant observations such as image geometry. |
| S3: Verify | Run and inspect a bounded trial before full execution; validate its output and address observed failures. | Pilot tool records and artifacts; checks of labels, coordinates, mask geometry, or the task’s analysis format. |
| S4: Execute | Complete the declared workload, save the required outputs, and retain their links to the inputs, method, and limitations. | Materialized artifacts, item-count checks, schema validation, and provenance linking full outputs to the inspected inputs and pilot. |
| S5: Submit | Validate the final artifact, submit one consistent result, and report only supported status and uncertainty. | Independent artifact checks, an accepted submission record, and a final response consistent with that record. |
| Setting | CoWork | MedResearch |
|---|---|---|
| Policy | Qwen3.8-27B ( Qwen Team, 2026c ) | |
| Training objective | Full-parameter GRPO | |
| Precision | BF16 | |
| Optimizer | Adam | |
| Batch | 12 prompts 8 rollouts | 8 prompts 8 rollouts |
| Context limit | 49,152 tokens | 262,144 tokens |
| Method | Budget | Configuration |
|---|---|---|
| SFT-200 | 200 updates | Batch 8; constant LR |
| GRPO / RaR | 500 updates | Shared RL settings; outcome / process rewards |
| OPD / ReflectRL | 500 updates | Shared settings in Table 12 |
| ADAS | 40 candidates | Frozen LLM; development-score selection |
| COS-PLAY | 200 decision updates | SFT-200; 30-skill bank; 28 selection / 172 action updates |
| Method | O | A | T |
|---|---|---|---|
| Base + VERA LLM + VERA Harness (Ours) | 31.0 ★ | 45.3 ★ | 16.7 |
| Base + ADAS harness ( Hu et al., 2025 ) | 20.7 ★ | 22.3 | 19.1 |
| Base + COS-PLAY skill bank ( Wu et al., 2026b ) | 19.9 | 18.2 | 21.6 ★ |
| Base + OPD ( Agarwal et al., 2024 ) | 19.8 | 15.6 | 24.0 ★ |
| Base + VERA LLM (harness frozen) | 18.9 | 23.6 ★ | 14.2 |
| Base + VERA Harness (LLM frozen) | 18.2 | 21.3 | 15.1 |
| Method | O | A | T |
|---|---|---|---|
| Base + VERA LLM + VERA Harness (Ours) | 69.1 ★ | 76.8 ★ | 61.3 ★ |
| Base + VERA Harness (LLM frozen) | 56.1 ★ | 63.4 ★ | 48.8 |
| Base + ADAS harness ( Hu et al., 2025 ) | 55.5 | 61.6 | 49.5 ★ |
| Base + VERA LLM (harness frozen) | 43.3 | 50.6 | 36.1 |
| Base + OPD ( Agarwal et al., 2024 ) | 41.2 | 51.8 | 30.6 |
| Base + GRPO ( Shao et al., 2024 ) | 36.3 | 47.0 | 25.7 |