cs.MASep 29, 2026

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

Authors: Zhibang Yang, Xinke Jiang, Yuxuan Liu, Mingyu Zhang, Zhixin Zhang, Zhengxing Song, Yue Fang, Guohong Qiu, +4 more

Organizations: National Engineering Research Center of Software Engineering, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Key Laboratory of High Confidence Software Technologies, Ministry of Education, Beijing, China · Center on Frontiers of Computing Studies, Peking University, Beijing, China · Peking University Information Technology Institute (Tianjin Binhai), Tianjin, China

Abstract

LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should change, or reuse execution histories under a fixed objective. We address this gap by formulating dynamic-requirement collaboration as joint requirement tracking and local update. We introduce GitHarness, a pluggable Git-style framework that organizes requirement states and their corresponding harness work states into a branchable version history. A trainable Git Agent resolves requirement changes and selects a semantically compatible historical state. A unified version interface then restores that state and creates a new branch, enabling the underlying harness to exclude obsolete information, inherit compatible work, and focus execution on affected parts. The Git Agent is trained through interface-level black-box reinforcement learning, with downstream harnesses and task-execution models kept fixed. We also construct MTAgentBench, a verifier-preserving benchmark covering mathematical reasoning, text-to-SQL, agentic search, software engineering, and research synthesis. Experiments demonstrate strong task performance alongside effective requirement tracking, preservation of valid work, and efficient execution.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

    Jul 14, 2026Ruhan Wang, Yucheng Shi, Zongxia Li +7Agent HarnessArtificial Intelligence Agents

  2. Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

    Jul 30, 2026Haomin Qi, Xingliang Wang, Xuanqi Gao +9Coding AgentsEnvironment