Organizations: National Engineering Research Center of Software Engineering, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Key Laboratory of High Confidence Software Technologies, Ministry of Education, Beijing, China · Center on Frontiers of Computing Studies, Peking University, Beijing, China · Peking University Information Technology Institute (Tianjin Binhai), Tianjin, China
LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should change, or reuse execution histories under a fixed objective. We address this gap by formulating dynamic-requirement collaboration as joint requirement tracking and local update. We introduce GitHarness, a pluggable Git-style framework that organizes requirement states and their corresponding harness work states into a branchable version history. A trainable Git Agent resolves requirement changes and selects a semantically compatible historical state. A unified version interface then restores that state and creates a new branch, enabling the underlying harness to exclude obsolete information, inherit compatible work, and focus execution on affected parts. The Git Agent is trained through interface-level black-box reinforcement learning, with downstream harnesses and task-execution models kept fixed. We also construct MTAgentBench, a verifier-preserving benchmark covering mathematical reasoning, text-to-SQL, agentic search, software engineering, and research synthesis. Experiments demonstrate strong task performance alongside effective requirement tracking, preservation of valid work, and efficient execution.
Figures & tables
Figure 1: When user feedback changes the umbrella color or removes a report section, reusing all prior work leaves stale content, while regenerating everything may alter unaffected parts. Selective updating limits changes to the affected work.
Figure 2: Framework of GitHarness .
Qwen3-32B
DeepSeek-V4-Flash
Harness
Cross-turn method
Math
SQL
Search
Code
Research
Math
SQL
Search
Code
Research
TCRAG
Native
72.0
53.0
1.0
3.0
25.02
74.0
69.0
34.0
56.0
42.54
+ Restart
73.0
64.0
2.0
5.0
16.67
88.0
66.0
41.0
61.0
43.45
+ U-Fold
67.0
38.0
7.0
9.0
14.46
68.0
71.0
14.0
59.0
45.72
+ GCC
71.0
63.0
6.0
11.0
15.66
55.0
53.0
42.0
53.0
32.37
+ GitHarness
89.0
70.0
8.0
19.0
38.45
93.0
80.0
56.0
62.0
46.65
Table 1: Final-task performance across harnesses. Scores use each benchmark’s native metric. Best results are shown in bold , and second-best results are underlined .
Figure 3: Performance versus token usage under StackPlanner with DeepSeek-V4-Flash. Bubble area denotes estimated API cost; lower right is better.
Math
SQL
Search
Code
Research
Component
Variant
Acc. ↑
Tok. ↓
EX ↑
Tok. ↓
EM ↑
Tok. ↓
Res. ↑
Tok. ↓
RACE ↑
Tok. ↓
Full GitHarness
94.0
60.0
78.0
96.1
58.0
653.9
74.0
1291.4
49.96
1327.8
Requirement Resolution
w/o Historical Version Selection
84.0
104.6
62.0
142.1
42.0
1892.6
44.0
4513.5
45.59
2299.6
w/o Inspect
82.0
97.6
70.0
123.4
44.0
1583.5
42.0
4215.4
43.50
2412.3
w/o Exact-State Reuse
88.0
107.8
76.0
130.3
54.0
963.3
63
2534.8
49.88
1360.6
Branch Execution
w/o ReuseAndPatch
92.0
65.4
75.0
104.3
53.0
862.4
69.0
1531.4
46.63
1597.0
Table 2: Inference-time ablations under StackPlanner with DeepSeek-V4-Flash. Each task reports its native performance metric and average total token usage (thousands).
Figure 4: Harness RL adaptation with fixed downstream execution: (a) final-task performance on the held-out test set, (b) training reward, and (c) reward-weight sensitivity on the validation set.
StackPlanner
OpenHands
Method
Math
SQL
Search
Code
Research
Math
SQL
Search
Code
Research
GitHarness
94.0
78.0
58.0
74.0
49.96
92.0
79.0
57.0
72.0
51.48
GitHarness -Skills
95.0 (+1.0)
83.0 (+5.0)
58.0 (0.0)
75.0 (+1.0)
52.03 (+2.07)
94.0 (+2.0)
81.0 (+2.0)
58.0 (+1.0)
74.0 (+2.0)
53.60 (+2.12)
Table 3: Effect of skill injection under DeepSeek-V4-Flash. Parentheses report changes from the corresponding skill-free GitHarness.
Math
Research
Method
H=10
H=15
H=20
H=25
H=30
H=10
H=15
H=20
H=25
H=30
Native
73.0
72.0
69.0
63.0
62.0
42.92
41.57
37.86
35.39
34.53
Restart
71.0
71.0
67.0
61.0
59.0
41.05
40.08
39.14
35.88
27.70
U-Fold
82.0
83.0
81.0
79.0
78.0
37.53
37.13
35.41
33.52
34.16
GitHarness
92.0
90.0
87.0
85.0
84.0
49.53
49.27
48.71
47.52
47.32
Table 4: Long-horizon performance under StackPlanner with DeepSeek-V4-Flash.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Change type
Operator
Effect
Completion
Reveal
Disclose previously unstated source-task information
OutputControl
Revise format, structure, length, or presentation
Refine
Request additional explanation, evidence, or detail
Elicitation
Constrain
Add a newly formed executable constraint
Comparison
Add an object, method, metric, or comparison dimension
Pivot
Introduce a related function while retaining compatible context
Appendix
Table 5: Requirement operators used to construct multi-turn interactions. Operator labels are hidden from evaluated agents.
Class
Generic operations
Purpose
Inspect
Glob , Grep , Read
Inspect files or recovered work units
Task
Domain tools
Perform domain-specific task operations
Edit
Write , Patch , Delete
Modify only the isolated candidate state
Finish
Finish
Return the result and request runtime finalization
Appendix
Table 6: Semantic action classes available to the update agent.
Dataset
RL Train
Validation
Test
Lang.
GSM8K
100
100
100
EN
BIRD
100
100
100
EN
BrowseComp-Plus
0
0
100
EN
SWE-bench Verified
0
0
100
EN
DeepResearch Bench
0
0
100
EN/ZH
Appendix
Table 7: Task-disjoint training, validation, and test subsets. Only the training trajectories are used for RL adaptation.
Domain
Model-visible tools and constraints
Persistent work products
Math
Deterministic calculator and scoped workspace read/write/patch operations
answer.md and calculation evidence
SQL
SQL generation from the public schema and BIRD evidence; no database access during inference
answer.sql
Search
Read-only BrowseComp retrieval returning top-5 passages with 512-token snippets
answer.json and runtime-recorded retrieval evidence
Research
Bounded read-only BoCha web search
report.md , report outline, and runtime-recorded retrieval evidence
Code
Repository inspection, editing, shell execution, and repository-local test execution in an isolated workspace; no access to hidden tests
Repository patch
Appendix
Table 8: Domain-specific tools and persistent work products.
Parameter
Value
Parameter
Value
Models and execution
Trainable policy
Qwen3-8B Git Agent
Downstream executor model
DeepSeek-V4-Flash
Query-equivalence judge
DeepSeek-V4-Flash
Adaptation
LoRA
Git-Agent tool steps
4
Downstream tool steps
16
Git-Agent context / response limit
16,384 / 4,096 tokens
Downstream response limit
4,096 tokens
Sampling
Appendix
Table 9: LoRA-GRPO training configuration for the Git Agent.
Figure 5: Performance versus token usage on Math and Code under StackPlanner with DeepSeek-V4-Flash. Bubble area denotes estimated API cost; lower right is better.
Figure 6: Performance versus estimated API cost across five domains under StackPlanner with DeepSeek-V4-Flash. Bubble area denotes average token usage; lower right is better.
Math
SQL
Search
Code
Research
Method
Tokens
Cost
Tokens
Cost
Tokens
Cost
Tokens
Cost
Tokens
Cost
Native
43.13
0.03
149.68
0.09
1,438.03
0.77
4,900.13
2.28
2,095.92
1.11
Restart
26.82
0.02
137.39
0.08
874.40
0.47
3,649.91
1.69
1,483.71
0.79
U-Fold
54.09
0.04
89.79
0.05
669.87
0.36
2,266.16
1.05
1,742.79
0.92
GCC
53.05
0.04
142.29
0.08
1,536.62
0.82
2,034.28
0.94
1,604.36
0.85
GitHarness
59.98
0.04
96.10
0.06
653.87
0.35
1,291.43
0.63
1,327.77
0.70
Appendix
Table 10: Detailed token usage (thousands) and API cost (USD) under StackPlanner with DeepSeek-V4-Flash.
Figure 7: Requirement-resolution actions across all domains and branch-execution strategies for SQL and Search.
Figure 8: Branch-execution strategies for Math, Code, and Research.
Method
Math
SQL
Search
Code
Research
GitHarness
93.0
80.0
56.0
62.0
46.65
GitHarness -Skills
96.0 (+3.0)
81.0 (+1.0)
55.0 (-1.0)
63.0 (+1.0)
50.92 (+4.27)
Appendix
Table 11: Effect of skill injection with TCRAG under DeepSeek-V4-Flash.
Domain
Additional semantic check
Math
Preserve the mathematical goal, givens, variables, units, relationships, active constraints, and requested output.
SQL
Preserve the database question, requested columns or aggregation, filters, grouping, ordering, limits, comparisons, and aliases. Treat fixed schema and SQL transport wrappers as harness context.
Search
Preserve the information need, entities, disambiguating clues, time range, requested answer granularity, and evidence or citation requirements.
Research
Preserve the research question, scope, subjects, time and region bounds, requested synthesis, deliverable structure, evidence, and citation requirements. Do not reduce the task to short-answer search.
Code
Preserve the software issue goal, repository-scoped entities, affected symbols or files, reproduction and failure behavior, expected behavior, version constraints, requested change scope, and active testing or output requirements.
Appendix
Table 12: Values used for the domain-specific semantic-check placeholder.
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behaviorally distributed, while modification requests describe what the system should do and repositories are organized by files and modules. Code search, repository indexing, and long-context processing ease inspection, but still leave this behavior-to-code mapping to be recovered by hand. Behavior localization is therefore a central bottleneck in harness evolution. We introduce the Harness Handbook, a behavior-centric representation synthesized automatically from a harness codebase via static analysis and LLM-assisted structuring, linking each behavior to its corresponding source. We also introduce Behavior-Guided Progressive Disclosure (BGPD), which guides agents from high-level behaviors to relevant implementation details and verifies candidate locations against the current source. On diverse modification requests from two open-source harnesses, Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens, with the largest gains on scattered sites, rarely executed paths, and cross-module interactions. Evolving complex agentic systems thus depends not only on generating edits, but also on determining where those edits should be made.
Ruhan Wang, Yucheng Shi, Zongxia Li +7
Tencent · Indiana University · University of Maryland, College Park +2
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validates the lifecycle from a healthy base to a task state and a restored state. By deriving multiple tasks grounded in developer evidence from maintained environments, Change2Task provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task construction effort. We evaluate the system through five common and widely adopted coding agent task families: Bug Fix, Feature Addition, Test Generation, Application Programming Interface Migration, and Security Repair. Starting from 1,130 source changes eligible for construction, Change2Task achieves 79.6% verified task construction success across these task families. On a matched candidate set, it recovers 29.2% more verified tasks than a construction baseline based on pull requests. Historical and reconstructed cases achieve up to 98.0% matched outcome agreement under agent evaluation, while reuse of modern bases reduces measured expenditure across the complete pipeline by 10.8%.
Haomin Qi, Xingliang Wang, Xuanqi Gao +9
1Microsoft · 3Zhejiang University · 4Xi’an Jiaotong University +1
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench2.1, and from 2.8% to 8.3% on OSWorld2.0. It also raises Claude Opus4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.