Chronos Enables Code Agents to Reason over Software Evolution
Organizations: Zhejiang University, China · National Industrial Information Security Development Research Center, China
Abstract
Historical pull requests record the design decisions, compatibility constraints, and implementation patterns behind a codebase's current state. Experience relevant to a new task can span related changes whose descriptions emphasize different concerns. We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents. Chronos distills merged pull requests into structured experience cards and connects them through a typed graph of code-level, developer-intent, and organizational relations. Semantic search identifies entry cards, and weighted multi-hop expansion retrieves connected changes for selective reading. The same memory guides candidate generation and patch selection: a patch-focused change agent and a validation-strategy agent each develop a patch, and an evolution steward consults history to select between them. On SWE-Bench Verified, the full workflow improves SWE-Agent across all six evaluated LLM backbones, raising the mean resolution rate from 69.2% to 72.9% and reaching 79.8% with MiniMax M2.5. With the same backbone, it raises resolution rates from 48.3% to 51.7% on SWE-Bench Pro and from 41.0% to 43.5% on FEA-Bench Lite. Both experience-guided single-agent variants also outperform the base agent. In a human evaluation on 100 tasks with ten cards retrieved per task, graph-grounded retrieval increases the mean number of useful cards from 1.24 to 2.87 over flat semantic retrieval. These results demonstrate the value of PR relations for retrieving useful repository experience and of the evaluated workflows for applying that experience during patch generation and selection.
Figures & tables
| Tier | Edge type (forward / reverse) | Criterion | Cov. | Weight |
| T1 | precedes / succeeds | File-level Jaccard overlap between PRs. | 62% | 1.0–1.5 |
| extends / extended_by | A later PR continues development on files newly introduced by an earlier PR. | 13% | 1.0 | |
| T2 | follows / followed_by | Same author, time gap 60 days, and directory Jaccard . | 73% | 0.70–1.05 |
| backport_of / backported_to | Backport cues detected from title, body, commit messages, or base branch. | 5% | 0.70 | |
| references_to / referenced_by | Explicit #N citation in PR body or commit messages. | 46% | 0.70 | |
| T3 | same_reviewers | Reviewer overlap (Jaccard ). | 28% | 0.30–0.45 |
| Candidate outputs | Selection action |
| Two identical, non-empty patches | Retain the shared patch directly. |
| One non-empty patch | Retain the available patch directly. |
| Two empty patches | Submit an empty patch. |
| Two distinct, non-empty patches | Randomize their X/Y order and invoke the steward. |
| Benchmark | Tasks | Repos. | Task emphasis |
| SWE-Bench Verified | 500 | 12 | Human-verified issue resolution |
| SWE-Bench Pro (public) | 731 | 11 | Longer-horizon repository-level changes |
| FEA-Bench Lite | 200 | 48 | New feature implementation |
| Benchmark | SWE-Agent | Chronos | |
| SWE-Bench Pro | 48.3 | 51.7 | |
| FEA-Bench Lite | 41.0 | 43.5 |
| Configuration | Rate (%) | |
| Chronos (SWE-Agent) | 79.8 | - |
| Chronos (mini-SWE-Agent) | 79.2 | - |
| w/o validation-strategy agent | 78.4 | |
| w/o patch-focused change agent | 77.8 | |
| w/o software-evolution graph | 79.0 | |
| SWE-Agent | 76.4 |
| Retrieval Method | Useful | Hit |
| Random | 0.18 | 14% |
| Embedding | 1.24 | 68% |
| Chronos | 2.87 | 89% |