Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agents, or external tools, while the model retains only the stale content in its context window. Existing agent runtimes provide little support for notifying the model that a previously observed fact has become stale, causing agents to reuse outdated observations and make incorrect claims about the current workspace state. We propose Concord, a context coherence framework that maintains the consistency between tool observation in agent context and the mutable sources from which they were derived. Concord links each observation to its source, detects source changes, and uses configurable handling policies to update, annotate, or suppress stale context before reuse. Concord is applicable across different agent runtimes and external resources, and can be easily extended to new runtime-resource settings. We implement Concord as a general framework, and instantiate a concrete use case to assess its effectiveness. We construct ConcordBench, where previously observed file contents become stale after subsequent edits. Across three evaluated frontier models, Concord produces answers consistent with the restored workspace state in all evaluated cases under these constructed conditions, matching the oracle on recover count for this benchmark, while using 46.4% fewer tokens than the strongest non-oracle baseline.
Figures & tables
Figure 1. A stale-observation failure in an agent-maintained workspace. After the agent records an old free-shipping threshold in its context, the configuration file is edited outside its view. The agent later answers from the stale observation, producing an answer that contradicts the current workspace state.
Figure 2. Architecture overview of Concord . A runtime- and source-agnostic Context Manager tracks tool observations as source-linked context blocks, connected to the agent runtime and to external resources through a Runtime Adapter and a Resource Adapter . Colored arrows denote the three data paths: block registration (green), source-change invalidation (red), and policy-driven context reconciliation (blue).
Figure 3. Example patch-and-restore race in ConcordBench-Pioneer . The A-side patch makes the agent observe mutated evidence, while the B-side restore reverts the workspace; the agent fails if the stale observation remains in its trajectory.
Recover Count ↑
Avg. Steps ↓
Tokens ↓
Claude Sonnet 5
Oracle †
40/40 (100%)
3.03
285,417
Naive
20/40 (50.0%)
4.40
589,720
Alert
26/40 (65.0%)
4.33
609,750
Diff
36/40 (90.0%)
7.10
1,079,420
Concord
40/40 (100%)
4.03
502,596
Table 1. Observed results on the 40-task HotpotQA-derived subset. Percentages are empirical recovery rates for this finite benchmark; † denotes the no-race oracle upper-bound setting.