The architecture of modern LLMs consists of a profound cognitive polarization. LLMs possess implicit intuition encoded in their parameters, yet rely on a disconnected, explicit mechanism to access the outside world. Agentic frameworks have not bridged this gap; instead, models are often compelled into pathological "induced amnesia." Under the prevailing "Retrieve-Always" paradigm, agents must distrust their internal knowledge, making every user interaction a "tabula rasa" event that must be checked externally. This creates reflexive dependence that can be thermodynamically wasteful, cognitively fragile, and susceptible to irrelevant context. We propose a return to first principles, operationalizing the biological maxim "Look Before You Leap." We introduce MARTA (Metacognitive Adaptive Retrieval and Thought Architecture), a neuro-symbolic framework that bridges parametric and non-parametric knowledge. Rather than treating retrieval as mandatory, MARTA models it as a cost, taking the leap only when perceived internal inadequacy warrants external information. By allowing the agent to gauge the entropy of its own thoughts before acting, MARTA enables deliberative retrieval and uncertainty-aware decision making. Our approach suggests that giving agents the capacity for introspection can restore a more efficient balance between internal knowledge and external information.
Figures & tables
Figure 1: Schematic of MARTA. The architecture fuses the Introspective State (Internal Entropy) with Extrospective Affordances (Memory Keys) via a Cross-Modal Attention Gate. This allows the agent to ”Look” (measure uncertainty) before it ”Leaps” (retrieves).
Control Topology
Single-Hop (Code)
Multi-Hop (Reasoning)
Latency Overhead
Static Projector
98.2%
34.5%
8 ms
Blind Ablation (Null-Affordance)
96.5%
61.2%
9 ms
Open-Loop RAG
100.0%
88.1%
450 ms
MARTA (Ours)
99.2%
68.4%
12 ms
Table 1: Arbitration Competence Analysis. We compare the ability to retrieve necessary context across task complexity. The Blind Ablation outperforms the Static baseline by detecting uncertainty, but fails to match MARTA in Multi-Hop scenarios because it lacks the ability to verify the utility of the retrieved memory.
Control Architecture
Acceptance (Gold)
Rejection (Hard Negative)
Discriminative Gap ( Δ )
Static Projector
98.5%
21.8%
+20.3
MARTA (Ours)
99.8%
87.6%
+87.4
Table 2: Adversarial Robustness Profile. The Rejection Rate measures how often the control layer correctly refuses to ingest a deceptive ”Hard Negative” document. Higher is better.
Figure 2: Adversarial Rejection Capabilities. The Static Projector (Red) retrieves ”Hard Negatives” nearly as often as correct documents. MARTA (Blue) exhibits a sharp drop in acceptance for distractors, effectively acting as a hallucination filter.
Figure 3: Regulatory Decision Matrix. The chart illustrates MARTA’s action distribution. It correctly avoids retrieval (Grey bar) for parametric sufficiency and adversarial noise, while triggering retrieval (Blue bar) for genuine epistemic gaps.
Figure 4: Architectural Generalization. MARTA (Blue) maintains high robustness scores across Qwen, Llama, and Mistral, whereas the baseline (Red) varies significantly. This confirms the neuro-symbolic head is model-agnostic.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Temporal Budget Analysis. Note the logarithmic scale. The MARTA gating mechanism introduces a deterministic, constant-time overhead of ≈12 ms. In contrast, the standard Retrieval Pipeline is a high-latency operation ( ≈450 ms) involving dense vector embedding and HNSW graph traversal. The asymmetry is stark: the computational cost of arbitration is orders of magnitude lower than the cost of execution .
Figure 6: VRAM Utilization Profile. The MARTA controller consumes only 48MB of memory, a negligible allocation ( <0.3% ) compared to the 16GB required by the frozen 7B backbone. This hyper-efficiency allows the controller to reside permanently in HBM without necessitating model quantization or sharding.
Figure 7: Quantization Stability Analysis. Arbitration accuracy remains statistically invariant ( >91% ) even when the controller is compressed to 4-bit precision (NF4). This suggests that the epistemic signal is a macroscopic feature of the logit distribution, robust to the microscopic noise introduced by low-precision arithmetic.
Figure 8: Learning Dynamics. MARTA achieves 90% arbitration fidelity with only 1,000 training examples. This rapid saturation indicates that the task of ”epistemic arbitration” operates on a significantly lower-dimensional manifold than ”general reasoning,” allowing for rapid adaptation to new domains.
Figure 9: Causal Intervention Test. Comparison of the Full MARTA against a ”Feature-Suppressed” variant (frozen uncertainty signal). The massive performance collapse ( Δ≈37% ) isolates the specific contribution of epistemic awareness. Without access to its own confusion, the model defaults to semantic heuristics, failing to reject adversarial distractors.
Figure 10: Component Importance Analysis. Removing the Entropy Signal (Red) causes a catastrophic degradation to 76.4%, reinforcing its dominance. However, removing the Hidden Context (Amber) also degrades performance to 81.2%, indicating that uncertainty must be grounded in semantic context—the model must know what it is confused about.
Figure 11: Thermodynamic Efficiency. The ”cloud” represents the stochastic variance of raw training steps. The CEA objective (Navy) converges to a significantly lower asymptotic loss than the Baseline (Red). This efficiency stems from the contrastive formulation, which explicitly pushes ”hard negatives” away in the manifold, providing a denser gradient signal than simple classification.
Figure 12: Sensitivity Analysis. We measure the KL-Divergence of the gating distribution when the input query is perturbed. The high sensitivity of the Arbitrator (0.61 vs 0.25) confirms that it acts as a dynamic metacognitive gate, adjusting its retrieval threshold based on the quality of the query specification.
Figure 13: Cross-Domain Generalization. The Baseline performance collapses on the Medical domain (58.1%), typically hallucinating relevance because medical jargon creates spurious vector similarities. In contrast, MARTA maintains high precision ( >86% ). This supports the hypothesis that MARTA has learned a domain-agnostic ”Signature of Ignorance” : the statistical texture of the logits when the model is ”confused” is topologically identical whether the topic is Python or Oncology.
Figure 14: Calibration Analysis. MARTA (Navy) achieves an Expected Calibration Error (ECE) of 0.04, significantly superior to the Baseline (0.18). Crucially, the curve exhibits a slight ”S-shape” deviation, remaining marginally under-confident at high probabilities ( >0.9 ). This is a desirable ”Safety-First” characteristic, contrasting with the dangerous over-confidence of the Baseline.
Figure 15: Robustness to Prompt Variation. MARTA exhibits a tight distribution of outcomes ( σ=0.06 ) compared to the volatile Baseline ( σ=0.15 ). This stability indicates that the internal epistemic signal provides a robust Thermodynamic Anchor . While the semantic embedding of the prompt shifts, the model’s internal uncertainty regarding the user’s query remains constant, stabilizing the decision boundary.
Figure 16: The Epistemic Horizon. MARTA maintains robustness for up to 3 reasoning hops. Beyond this ”Effective Limit,” performance degrades non-linearly. This sharp drop-off suggests that for highly complex chains (4+ hops), the accumulation of aleatoric noise drowns out the epistemic signal. This validates our architectural choice: MARTA is an Arbitrator , not a Reasoner . For deep chains, it serves as the initial gatekeeper, but subsequent hops require iterative agentic re-verification.
Query Archetype
Input Query ( x )
Baseline Action
Entropy ( H )
Margin ( Δ )
MARTA Action
Parametric Fact
”What is the capital city of France?”
Retrieve (Redundant)
0.12 (Low)
0.95 (High)
Implicit ( ∅ )
Long-Tail Fact
”Details of the ISO-8601 date format spec.”
Retrieve (Correct)
0.84 (High)
0.15 (Low)
Explicit
Ambiguous
”Explain the function defined above.”
Retrieve (Hallucination)
0.65 (Med)
0.30 (Med)
Implicit ( ∅ )
Appendix
Table 3: Qualitative Arbitration Analysis. We display the Epistemic Signature and the resulting action. MARTA correctly identifies that ”Capital of France” requires no retrieval (High Margin, Low Entropy), whereas it actively retrieves for the specific ”ISO-8601” query. Crucially, in the Ambiguous case, the Baseline hallucinates a context, whereas MARTA detects the lack of specific affordance and correctly chooses to bypass.
Figure 17: Logic Puzzle Bypass Rate. The Baseline frequently retrieves irrelevant facts (34.2% bypass), confusing the complexity of the logic puzzle with a need for external facts. MARTA, detecting that the reasoning task is self-contained (the entities A, B, C do not exist in the corpus), correctly bypasses retrieval 88.5% of the time. This demonstrates a capability to identify when memory is contextually useless.
Figure 18: Safety Analysis. The histogram shows a clear separation: hallucinations (Red) generally exhibit higher entropy than correct facts (Navy). However, the overlap region (”Danger Zone”) represents False Confidence —misconceptions deeply embedded in pre-training where the model is confident but incorrect. This remains an open challenge for all uncertainty-based methods.
Figure 19: Robustness to Memory Compression. When raw chunks are replaced with LLM-generated summaries, the Baseline’s accuracy drops significantly (85% → 72%) due to the loss of exact lexical overlap. MARTA maintains robust performance (89.8%), confirming that the Cross-Modal Attention aligns with the semantic affordance of the memory (the ”gist”), not just surface-level keywords.
Predictive Performance
Computational Efficiency
Methodology
Control Mechanism
Recall@5
QA Acc.
Hallucination
Latency
Throughput
VRAM
Token Cost
( ↑ )
(F1, ↑ )
(Rate, ↓ )
(ms, ↓ )
(QPS, ↑ )
(GB, ↓ )
(Overhead, ↓ )
Naive RAG
Dense Retrieval
76.5
42.5
18.4%
450
22
14.2
N/A
Self-RAG
Reflection Tokens
81.2
58.4
9.1%
820
12
14.8
+24.5%
CRAG
Corrective Evaluator
82.5
61.2
7.5%
950
9
15.1
+31.2%
MARTA (Ours)
Epistemic Gating
83.1
63.8
6.2%
291
34
14.25
0.0%
Appendix
Table 4: Comprehensive System Performance Profile. We compare MARTA against standard and active RAG baselines using a Mistral-7B backbone on the PopQA dataset. Bold indicates the optimal result. MARTA dominates the Efficiency Landscape : it matches the high QA Accuracy (63.8%) of CRAG while maintaining the Throughput (34 QPS) of a sparse system. Crucially, it reduces the Hallucination Rate to 6.2% by effectively filtering adversarial distractors, achieving a superior safety profile without the 30%+ token overhead incurred by generation-based controllers.
Figure 20: Efficiency-Performance Analysis. Bars represent F1 Score (Left Axis), while the red line represents Inference Latency (Right Axis). While Self-RAG and CRAG achieve high accuracy, they incur a massive latency penalty ( >800 ms) due to the serialization of the reasoning process. MARTA matches their accuracy (63.8%) but maintains the latency profile of a sparse system (291ms), effectively defining the optimal operational boundary.
Figure 21: Zero-Regression Verification. The performance of the backbone remains statistically identical with and without MARTA attached. This confirms that the module is Non-Invasive : when the epistemic signal is low (as in standard math or coding tasks), the controller output probability vanishes, preserving the original reasoning pathways of the host model.
Figure 22: Cross-Domain Robustness Analysis. MARTA consistently outperforms the Static RAG baseline across all modalities. The gain is most pronounced in specialized verticals like Finance (+26.7%) and Medicine (+26.3%) . This validates the hypothesis that MARTA effectively detects ”Out-of-Distribution” concepts: a generalist model exhibits distinct thermodynamic signatures when facing specialized jargon, correctly triggering the arbitrator to fetch definitions.
Archetype
User Query
Baseline (Naive RAG) Output
MARTA (Ours) Output
1. Fabrication (Hallucination Test)
”Summarize the Treaty of Westphalia II signed in 1995.”
Retrieves irrelevant history docs. ”The Treaty of Westphalia II was a peace agreement signed to end the conflict in the Balkans , reaffirming the principles of sovereignty…” [Result: Confident Hallucination]
Trigger: HHigh , Affordance ≈0 Action: Implicit ( ∅ ) ”I cannot find any record of a ’Treaty of Westphalia II’ signed in 1995. The original Peace of Westphalia was signed in 1648.” [Result: Safe Correction]
2. Precision (Medical Recall)
”What is the interaction between Warfarin and Vitamin K?”
Retrieves general nutrition blogs. ”Vitamin K is essential for health and can be taken freely with Warfarin to improve bone density.” [Result: Dangerous Error]
Trigger: HHigh , Affordance ≈1 Action: Explicit Retrieval ”WARNING: Vitamin K decreases the effectiveness of Warfarin , increasing the risk of blood clots. Consistent intake is required.” [Result: Grounded Safety]
3. Reasoning (Distractor Test)
”A is taller than B. B is taller than C. Who is shortest?”
Retrieves random height stats. ”Based on average height data, B is usually around 5’9” , making them shorter than A…” [Result: Context Distraction]
Trigger: HLow , Logic Pattern Action: Implicit ( ∅ ) ” C is the shortest. This is a transitive deduction: if A>B and B>C , then C must be the minimum.” [Result: Focused Reasoning]
Appendix
Table 5: Comparative Generation Analysis. We highlight critical divergences in model behavior. Red text indicates hallucinations or dangerous misinformation. Green text indicates correct, grounded generation or appropriate refusal. MARTA demonstrates superior semantic discipline, refusing to hallucinate when context is missing and retrieving with high precision when necessary.
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
University of California Los Angeles · University of Washington · Northwestern University
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting self-awareness capability, the ability to discern whether a problem requires necessary external resources or can be solved via internal parametric knowledge. To address this, we introduce KAPRO (Knowing-Acting Quadrant PRObe), a framework that evaluates cognitive-behavioral alignment by decoupling an agent's metacognitive judgment (Knowing) from its spontaneous execution (Acting). We further construct KAware, a dataset rigorously partitioning tasks into external, internal, and hybrid subspaces to systematically probe these epistemic boundaries. Extensive experiments across diverse agent architectures show that self-awareness capability is strongly correlated with task success but degrades sharply in internal-capability settings. Moreover, open-source and instruction-following models exhibit stronger tool overuse due to shallow pattern matching, while proprietary and reasoning-oriented models demonstrate more reliable cognitive gating. Benchmark and codes are available at https://github.com/AI-Santiago/KAware.
Yifan Li, Shengbin Yue, Boyu Feng +6
The Chinese University of Hong Kong · Fudan University · Tencent +2
Large Language Model (LLM) agents are deployed in complex environments -- such as massive codebases, enterprise databases, and conversational histories -- where the relevant state far exceeds their context windows. To navigate these spaces, an agent must iteratively explore the environment to find relevant information. However, without explicit infrastructure, an agent's working memory can degrade into lossy representations of the search state, resulting in redundant work (e.g. repetitive looping) and premature stopping. In this work, we formalize this challenge as the Context Gathering Decision Process (CGDP), a specialized Partially Observable Markov Decision Process, where an agent's objective is to adaptively refine its belief state to isolate the necessary information for a task. We model an LLM's behavior as approximate Thompson Sampling within this CGDP, and introduce a predicate-based method that decomposes an LLM's implicit search into explicit and modular operations. We then derive two plug-and-play interventions for iterative LLM agents: a persistent, predicate-based belief state that bounds context while preserving multi-hop reasoning, and a programmatic exhaustion gate that halts unproductive search without premature stopping. Across four methods and three question-answering domains, we empirically validate that replacing an LLM's implicit state with our CGDP-motivated belief state improves multi-hop reasoning by up to 11.4%; while the modular programmatic exhaustion detection saves up to 39% of tokens without any degradation in agent performance. Ultimately, we argue that framing the LLM agent loop as a CGDP can guide the design of modular, non-interfering improvements to agentic search harnesses.