Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
Organizations: Independent Researcher · SAGE7 AI, Georgetown, Texas, USA
Abstract
Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic. Existing agent-memory systems (e.g., MemGPT, Zep, A-MEM) gate retrieval by content, ownership, and role, not derivation, missing a cached insight that embeds a forbidden column. We introduce the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result, gated by a retrieval policy that serves a hit only when the requester is authorised for every column touched. Provided lineage recording is complete, we prove by construction that the policy blocks retrieval of results derived from a sensitive column outside the requester's permissions, at O(n) worst case -- a conditional design guarantee, not an empirical claim, that excludes derived features encoding sensitive information without naming their source. Eliminating measured leakage required 75-90% recorded lineage completeness, so we treat 90% as a conservative deployment target. Across six experiments, lineage-gated retrieval removes the 18.8-25.5% cross-department leakage naive content-gated memory suffers, keeping 81.5-82.6% of memory reuse at 13.8 microsecond worst-case overhead. A real-agent proof-of-concept with LLM-generated SQL is consistent with the guarantee: zero leaks over 9 round-trips, two conflicts caught automatically -- though a feasibility demonstration, not evidence of production viability. This offers a practical governance layer for shared agent memory, complementing source-layer access control and supporting EU AI Act compliance.
Figures & tables
| System | Content tag | Role ACL | Column lineage | Drift detect | Retrieval gate |
|---|---|---|---|---|---|
| MemGPT | ✗ | ✗ | ✗ | ✗ | ✗ |
| Zep | ✓ | ✗ | ✗ | ✗ | ✗ |
| A-MEM | ✓ | ✗ | ✗ | ✗ | ✗ |
| Governed Memory | ✓ | ✓ | ✗ | ✗ | ✓ |
| SSGM | ✓ | ~ | ✗ | ~ | ✓ |
| Collaborative Memory | ~ | ✓ | ✗ | ✗ | ✓ |
| Field | Description |
|---|---|
| metric_name | Business KPI name ( e.g. , churn_rate ) |
| value , owner_dept , epoch | Result, producing department, TTL time index |
| lineage | Derivation graph: sequence of steps |
| sensitivity_tags | (Eq. 1 ), cached |
| definition_hash | (Eq. 2 ), cached |
| Operation | Cplx. | ||
|---|---|---|---|
| Gate-check predicate | s | s | |
| Write — conflict detection | s | s | |
| Retrieve — worst case | s | s | |
| Retrieve — expected case | s | s |
| ID | Threat | Status | Production control |
|---|---|---|---|
| T1 | Sensitive-column leakage via cached result | Addressed | Lineage gate (Alg. ) |
| T2 | Silent metric-definition drift | Addressed | Conflict detector (Alg. ) |
| T3 | Stale sensitive data after policy update | Partial | TTL + lazy recompute (Cor. 3 ) |
| T4 | Inference attacks (undeclared derived features) | No | Manual review of feature engineering; out of scope |
| T5 | Malicious/incomplete lineage fabrication | Partial | Automatic extraction ( SQLGlot /LINEAGEX), § VI-B |
| T6 | Aggregate inference across retrievals | No | Differential privacy layer (future work) |
| Dept. | Permitted (summary) | Blocked (in ) |
|---|---|---|
| Finance | All columns | None |
| Marketing | Commercial, order data | ssn, income, email |
| Support | Customer, ticket meta | ssn, income, email |
| Synthetic (5-tab.) | TPC-H (8-tab.) | |||
| Metric | Naive | LA | Naive | LA |
| Leak rate (%) | ||||
| Reuse rate (%) | ||||
| Conflict recall (%) | ||||
| † Formal guarantee (Theorem 1 ), not an empirical discovery. | ||||
| Detector | Prec. | Recall | F 1 | LO correct |
|---|---|---|---|---|
| D1 — Exact hash | 0.651 | 1.000 | 0.789 | 8/8 |
| D2 — Jaccard ( ) | 0.676 | 0.821 | 0.742 | 3/8 |
| D3 — Column-graph | 1.000 | 0.714 | 0.833 | 0/8 |
| Venkata Sangaraju is an independent researcher working on memory and data-governance architectures for enterprise artificial intelligence (AI) agents. His research interests include lineage and provenance tracking for AI agent systems, privacy-preserving retrieval, column-level access control, and the application of formal safety guarantees to shared multi-agent memory infrastructure. He is the corresponding author of this work (ORCID: 0009-0001-7716-1342). |
| Sudhir Vissa is with SAGE7 AI, Georgetown, Texas, USA. His research interests include enterprise AI agent architectures, data governance, and applied infrastructure for multi-agent deployments, including the metric-definition and access-control problems addressed in this work (ORCID: 0009-0003-7865-0863). |