Knowledge-intensive language-model systems typically represent external knowledge as text chunks or static graphs, with limited support for concept evolution, point-in-time reasoning, and distinctions between validated and inferred knowledge. We introduce the Concept Lifecycle Model (CLM), which represents concepts as persistent, graph-grounded, temporally versioned entities with explicit provenance and epistemic status, and Concept-Grounded Attention (CGA), which injects concept-graph structure into transformer computation through graph-biased self-attention (Form A) and gated cross-attention over concept nodes (Form B). We evaluate the framework in controlled settings using disabled-mechanism baselines. On 200 MuSiQue and HotpotQA questions with retrieval fixed, concept-graph retrieval recovers explicit multi-hop paths but does not improve evidence recall. Form A appears to steer attention, with 2.76 times more attention on gold than distractor concepts, but the same ratio occurs when Form A is disabled; the learned bias is negligible and no answers change. An identity-preserving Form B improves F1 from 0.188 to 0.221, but control concepts yield 0.213, indicating that most of the gain reflects added capacity. On LongMemEval, explicit temporal representation improves answer accuracy by 13 to 25 points across all tested generators, up to 122B parameters, while simplified CLM version resolution performs similarly to dated serialization because concept identity is not established reliably. On a synthetic source-independence task, protocol-derived epistemic status reduces unsupported assertions from 28% to 0.1% in a fine-tuned small model and from 19-68% to 0-5% in 72-122B models. Overall, the results support making temporal validity and epistemic status explicit, while showing that graph-attention diagnostics are not informative without disabled-mechanism controls.
Figures & tables
System
EM
F1
Passage R@5
Closed book
0.100
0.145
–
Dense RAG
0.110
0.164
0.687
Hybrid graph retrieval
0.130
0.188
0.662
Serialised concept graph
0.135
0.178
0.662
CGA-Form A
0.130
0.188
0.662
CGA-Form B (initial)
0.000
0.006
0.662
Table 1: Answer-level results on the 200 multi-hop test questions (flan-t5-small, greedy decoding). Passage Recall@5 refers to the retrieved evidence
Figure 1: Gold-to-distractor attention ratio for 197 test questions with Form A disabled ( θ=0 ) and with the trained Form A coefficients. Left: distributions. Right: per-question comparison; all points lie on the diagonal. The apparent evidence of graph-steered attention is reproduced exactly without Form A
Scale
Mean ∣BG∣
∣Δℓ∣
ρ
Changed
F1
Ratio
0
0
0
0
–
0.188
2.759
1
0.002
0.006
0.0003
0
0.188
2.758
10
0.020
0.063
0.003
2
0.188
2.748
100
0.196
0.654
0.033
9
0.190
2.648
1,000
1.957
7.38
0.338
122
0.157
2.257
10,000
19.57
10.01
0.424
136
0.111
2.445
Table 2: Effect of scaling the trained Form A coefficients. ∣Δℓ∣ : median maximum change in decoder logits; ρ : relative change of the encoder output; changed: answers differing from θ=0
Variant
Train
F1
Δ F1
ρ
Gate
Gold/distr.
No Form B
–
0.188
–
–
–
–
Initial
0
0.003
−0.185
0.968
0.50
1.15
Initial, η=0
0
0.003
−0.185
0.968
0.00
1.15
Initial
60
0.005
−0.183
0.968
0.62
2.47
Identity-preserving
0
0.188
0.000
0.000
0.02
1.19
Identity-preserving
60
0.188
0.000
0.000
0.34
2.53
Table 3: Form B diagnosis and repair (mean over seeds 13, 42, 77; untrained rows are seed-independent in F1). ρ : relative change of the hidden state at the Form B layer. Gold/distr.: mean Form B attention per gold and per distractor concept
Configuration
F1
Δ
95% CI
Form B, real concepts, 1 epoch
0.195
+0.008
[−0.002,0.021]
Form B, mismatched concepts, 1 epoch
0.195
+0.008
[−0.002,0.021]
Form B, real concepts, 3 epochs
0.221
+0.033
[0.007,0.062]
Form B, mismatched concepts, 3 epochs
0.213
+0.025
[−0.003,0.056]
Query-conditioned Form A
0.197
+0.010
[0.000,0.023]
Query-conditioned Form A, shuffled
0.197
+0.010
[0.000,0.023]
Table 4: Controls for added trainable capacity (500 training questions, mean over seeds 13, 42, 77). Δ : paired F1 difference to the same model without the mechanism (flan-t5-small: F1 0.188), with 95% bootstrap CI
System
Mean
P50
P95
P99
Hybrid graph retrieval
133.9
122.6
212.8
293.6
CGA-Form A
155.3
134.1
279.9
451.0
Table 5: Controlled generation latency in milliseconds (CPU, float32)
System
Knowledge update
Temporal reasoning
All
( n=78 )
( n=133 )
( n=211 )
Dense (no dates)
0.410
0.211
0.284
Dense, dated
0.551
0.316
0.403
Recency only
0.603
0.263
0.389
CLM resolution
0.526
0.293
0.379
Table 6: Temporal resolution on LongMemEval (oracle sessions, flan-t5-large, k=4 retrieved user turns). Answer accuracy: normalised gold answer contained in the prediction
Raw text
Graph, no status
CLM status
Accuracy
0.802
0.809
0.999
validated
0.853
0.757
1.000
hypothesised
0.157
0.293
1.000
competing
1.000
1.000
0.997
invalidated
1.000
1.000
1.000
promoted
1.000
0.997
1.000
Table 7: Epistemic evaluation on the synthetic graph (500 test instances, 100 per category, disjoint entity names; flan-t5-small fine-tuned on 2,000 instances; mean over three seeds). Unsupported: presenting a hypothesised claim as fact, relying on a retracted claim, or preferring an unconfirmed competitor
Generator
No dates
Dated
Recency only
CLM
flan-t5-large (780M)
0.251
0.379
0.370
0.365
qwen2.5-72b
0.422
0.592
0.436
0.607
gpt-oss-120b
0.408
0.659
0.445
0.678
qwen35-122b
0.408
0.611
0.408
0.597
Table 8: Temporal resolution on LongMemEval across generators (211 questions, LLM-judged accuracy; k=4 retrieved user turns, identical evidence for all systems)
Model
Measure
Raw text
Graph, no status
CLM status
qwen2.5-72b
Accuracy
0.798
0.566
0.986
Unsupported
0.320
0.677
0.000
gpt-oss-120b
Accuracy
0.770
0.566
0.864
Unsupported
0.190
0.657
0.003
qwen35-122b
Accuracy
0.586
0.590
0.920
Unsupported
0.537
0.670
0.050
Table 9: Epistemic evaluation with large models (500 test instances; no fine-tuning; the confirmation rule is stated in the instruction for every input format). Unsupported: presenting a hypothesised claim as fact, relying on a retracted claim, or preferring an unconfirmed competitor
Large Language Models (LLMs) struggle to incorporate new knowledge without forgetting or costly retraining. We propose DYNA, a lightweight framework that augments a frozen LLM with a temporal knowledge graph where events are nodes and temporal relations are directed, timestamped edges. The graph serves as an external, updatable memory. At query time, DYNA retrieves relevant nodes via random walks and centrality measures, then augments the LLM's response. Evaluated on three temporal recall tasks, DYNA reduces catastrophic forgetting by ~7% compared to fine-tuning and improves temporal ordering by ~5% over standard RAG. Higher graph clustering coefficients correlate with better retrieval, showing that graph structure matters. Contributions: (1) episodic memory as temporal KG, (2) retraining-free LLM augmentation, (3) graph properties as predictors of retrieval performance.
Ali Sarabadani, Mahtab Tajvidiyan
Department of Computer Engineering and Information Technology, University of Qom, Qom, Iran
Large language model assistants are increasingly expected to retain and reason over information accumulated across many sessions. We introduce EngramaBench, a benchmark for long-term conversational memory built around five personas, one hundred multi-session conversations, and one hundred fifty queries spanning factual recall, cross-space integration, temporal reasoning, adversarial abstention, and emergent synthesis. We evaluate Engrama, a graph-structured memory system, against GPT-4o full-context prompting and Mem0, an open-source vector-retrieval memory system. All three use the same answering model (GPT-4o), isolating the effect of memory architecture. GPT-4o full-context achieves the highest composite score (0.6186), while Engrama scores 0.5367 globally but is the only system to score higher than full-context prompting on cross-space reasoning (0.6532 vs. 0.6291, n=30). Mem0 is cheapest but substantially weaker (0.4809). Ablations reveal that the components driving Engrama's cross-space advantage trade off against global composite score, exposing a systems-level tension between structured memory specialization and aggregate optimization.
LLM-based knowledge-graph question answering (KGQA) delegates graph traversal to language models, turning each question into a sequence of local relation-selection decisions repeated across beams and hops. A common but untested default is to serialize the complete partial path into every routing prompt, even though the controller already maintains this path as exact symbolic state. Bounded Path Context (BPC) decouples these two roles: the controller retains full paths in symbolic memory for answer extraction and audit, while the relation-selection prompt exposes only the question, the current entity, outgoing relation candidates, and at most the last K hops. A controlled sweep over K -- fixing graph neighborhoods, beam budget, depth, decoding, and answer-extraction format -- shows that bounded histories match or exceed full-history prompting on complete WebQSP and CWQ test sets with Qwen3.5-9B-AWQ: K=1 achieves 0.487 answer-set F1 on WebQSP versus 0.472 for full history, and K=0 reaches 0.287 on CWQ versus 0.274, with 9.7% and 12.1% fewer input tokens respectively. At the 4B scale, K=1 remains the strongest setting on both benchmarks. Per-example analysis reveals that 71-84% of examples are unaffected by history length, while the affected cases expose when prior hops disambiguate versus distract. These results suggest that path serialization length is better treated as a tunable interface variable than as a default assumption in LLM-based graph controllers.
Xihang Shan, Ye Luo
School of Mathematical Sciences Xiamen University · School of Informatics Xiamen University