Just-In-Time Agent Memory with Runtime Agentic Research
Organizations: Beijing Academy of Artificial Intelligence · Peking University · Hong Kong Polytechnic University
Abstract
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To address this limitation, we propose Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime. A Memorizer preserves complete raw histories in a hierarchical page-store with compact navigational summaries, while a Researcher iteratively retrieves, inspects, and integrates evidence for each request. To train these memory-use behaviors, we introduce Memory-Gym, an evidence-grounded data synthesis pipeline covering nine task types across six domains, and optimize the Researcher through verified-trajectory supervised fine-tuning followed by Hint-guided Group Relative Policy Optimization. We demonstrate the effectiveness of JAM across a variety of benchmarks on agent memory and long-context processing, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches. To support reproducibility and future research, we release our anonymized source code at https://github.com/VectorSpaceLab/general-agentic-memory.
Figures & tables
| Method | Base LLM | LoCoMo | LME | NAQA | HotpotQA | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| SH | MH | TE | OD | OA | 56k | 112k | 224k | ||||
| Training-free | |||||||||||
| VANILLA | Qwen3.5-4B | 48.15 | 32.98 | 42.66 | 19.63 | 42.45 | 54.80 | 31.23 | 63.56 | 53.04 | 36.97 |
| RAG | Qwen3.5-4B | 49.04 | 34.39 | 44.11 | 17.68 | 43.37 | 56.20 | 32.86 | 49.97 | 44.83 | 49.87 |
| A-MEM | Qwen3.5-4B | 46.99 | 29.88 | 42.56 | 18.65 | 41.17 | 56.80 | 31.30 | 27.08 | 25.39 | 28.32 |
| Mem0 | Qwen3.5-4B | 34.84 | 26.45 | 34.47 | 16.83 | 32.10 | 34.40 | 26.72 | 30.19 | 27.57 | 24.53 |
| Training data | LoCoMo | LME | NAQA | HotpotQA |
|---|---|---|---|---|
| w/o SFT | 48.59 | 48.80 | 34.98 | 45.44 |
| MemAgent data | 48.60 | 55.20 | 31.98 | 54.53 |
| Memory-R1 data | 48.90 | 57.20 | 32.42 | 49.81 |
| SH only | 50.59 | 58.80 | 34.40 | 50.45 |
| MH only | 48.97 | 57.60 | 37.42 | 51.96 |
| SU only | 49.43 | 59.80 | 40.94 | 50.86 |
| Researcher | 64K | 128K | 256K | Overall |
|---|---|---|---|---|
| Untrained | 56.58 | 55.43 | 49.23 | 54.08 |
| SFT only | 68.42 | 64.13 | 64.62 | 65.67 |
| SFT+RL | 76.32 | 76.09 | 73.85 | 75.54 |
| Method | LoCoMo | LME | NAQA | HotpotQA |
|---|---|---|---|---|
| w/o Training | 48.59 | 48.80 | 34.98 | 45.44 |
| SFT Only | 50.70 | 60.20 | 41.59 | 52.08 |
| GRPO w/o Hint | 51.42 | 62.40 | 43.62 | 55.64 |
| w/o Memorizer | 49.06 | 64.00 | 38.87 | 57.72 |
| w/o Researcher | 43.30 | 58.40 | 32.51 | 39.10 |
| w/o BM25 | 45.90 | 62.40 | 38.88 | 52.14 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Evaluator | Accept rate | Agreement | Precision | Recall | F1 | |
|---|---|---|---|---|---|---|
| Human reference | 62.0 | – | – | – | – | – |
| MiniMax-2.5 | 58.5 | 91.5 | 0.823 | 95.7 | 90.3 | 92.9 |
| GPT-5.5 | 60.0 | 92.0 | 0.832 | 95.0 | 91.9 | 93.4 |
| Qwen3.5-122B-A10B | 57.5 | 90.5 | 0.803 | 95.7 | 88.7 | 92.1 |
| Quality criterion | Applicable | Passed | Pass rate (%) | Agreement (%) | |
|---|---|---|---|---|---|
| Evidence support | 200 | 196 | 98.0 | 99.5 | 0.886 |
| Answer uniqueness | 134 | 131 | 97.8 | 99.3 | 0.853 |
| Leakage-free | 200 | 199 | 99.5 | 100.0 | 1.000 |
| Task consistency | 200 | 197 | 98.5 | 99.5 | 0.855 |
| Coverage | 66 | 64 | 97.0 | 98.5 | 0.792 |
| Completeness | 66 | 63 | 95.5 | 98.5 | 0.849 |
| Outcome / primary error type | Count | Percentage (%) |
|---|---|---|
| Evidence-grounding failure | 33 | 33.0 |
| Ambiguity/non-uniqueness | 22 | 22.0 |
| Query leakage or construction issue | 16 | 16.0 |
| Task-consistency failure | 14 | 14.0 |
| Coverage/completeness failure | 11 | 11.0 |
| Correctly rejected (subtotal) | 96 | 96.0 |
| Setting | Offline Tokens | Offline Time | Online Tokens | Online Time | Avg. Rounds | F1 |
|---|---|---|---|---|---|---|
| (/workspace) | (s/workspace) | (/query) | (s/query) | |||
| Flat Raw Store | 0 | 0 | 8174.29 | 17.11 | 3.57 | 49.06 |
| JAM Workspace | 36924.51 | 79.53 | 6720.61 | 13.81 | 3.16 | 52.09 |
| Benchmark | Stopped within 5 rounds (%) | Mean | Median | P90 | Budget exhausted (%) |
|---|---|---|---|---|---|
| LoCoMo | 95.58 | 3.16 | 3 | 4 | 0.00 |
| LongMemEval | 82.80 | 3.82 | 3 | 6 | 0.40 |
| NarrativeQA | 84.00 | 3.92 | 3 | 8 | 0.00 |
| HotpotQA | 43.83 | 7.11 | 6 | 14 | 4.70 |
| Benchmark | Run 1 | Run 2 | Run 3 | Mean | Std. |
|---|---|---|---|---|---|
| LoCoMo F1 | 52.09 | 52.68 | 51.10 | 51.96 | 0.80 |
| LongMemEval Acc. | 65.20 | 64.60 | 65.00 | 64.93 | 0.31 |
| NarrativeQA F1 | 45.72 | 43.37 | 45.21 | 44.77 | 1.24 |
| HotpotQA F1 | 60.84 | 58.63 | 58.42 | 59.30 | 1.34 |
| Method | LoCoMo | NarrativeQA | HotpotQA |
|---|---|---|---|
| Evaluated with our judge | |||
| RAG | 58.25 | 48.00 | 53.13 |
| A-MEM | 54.74 | 45.00 | 35.68 |
| Mem0 | 46.23 | 43.00 | 38.02 |
| MemoryOS | 47.66 | 40.00 | 30.21 |
| LightMem | 48.57 | 40.00 | 41.15 |
| Memorizer backbone | Build time (s/workspace) | Online time (s/query) | Avg. rounds | F1 |
|---|---|---|---|---|
| Qwen3.5-4B (default) | 79.53 | 13.81 | 3.16 | 52.09 |
| Qwen3.5-122B-A10B | 132.43 | 14.36 | 3.25 | 51.30 |
| GPT-5.5 | 117.65 | 13.19 | 3.03 | 52.62 |
| Researcher Prompt Template |
|---|
| Your Role. You are a Research Agent exploring a hierarchical knowledge base to answer a question. Knowledge Base Structure. A Knowledge Base Overview is provided at the end of this system prompt. It includes a summary of the knowledge base content and the full directory structure. Each folder has a README that summarizes the files and subfolders inside it. You can use these README files to quickly understand the workspace structure and decide which files to inspect. Available Tools. 1. search . Search for files matching a query in the knowledge base. It returns relevant file names, paths, and summaries. Use this tool to find potentially relevant files. Parameter: query , a keyword or phrase to search for in file content. 2. browse . Open a file and extract information relevant to a query. An AI assistant reads the file and returns a summary of query-relevant content. Parameters: path , the file path to browse; query , a specific query asking for the exact information needed from the file. 3. open . Open a folder to view its README summary and directory listing. Use this tool to understand the contents and structure of a directory. Parameter: path , the folder path to open. Process. Follow the think tool-use loop. (1) Think about what is known, what is missing, and which tool should be used. (2) Use one or more tools to gather information, with at most five tool calls per round. (3) Receive observations from the tools. (4) Think again based on the observations and decide the next step. (5) Repeat until enough information has been collected. (6) Output the final answer in <answer> tags. Output Format During Exploration. Each exploration round should contain one <think> block followed by one or more <tool_use> blocks. Each tool call must be wrapped in its own <tool_use> block and must contain valid JSON. <think> [Reason about what you know, what information is missing, which tool to use, and why.] </think> <tool_use> {"tool": "search", "query": "your search query"} </tool_use> <tool_use> {"tool": "browse", "path": "your browse path", "query": "your browse query"} </tool_use> Final Answer Format. When enough information has been collected, output a final answer in the following format and stop the exploration loop. <think> [Summarize the collected evidence and reasoning.] </think> <answer> {"answer": "Your comprehensive answer to the question", "sources": ["/path/to/source1.md", "/path/to/source2.md"], "notes": "Additional notes or caveats"} </answer> Guidelines. First check the Knowledge Base Overview, since it already provides the directory tree. Use multiple tools in one round when this can speed up exploration, but use at most five tool calls per round. Use multiple rounds of thinking and tool use until the evidence is sufficient. When outputting <answer> , make sure the answer is grounded in collected evidence and includes source paths. Each <tool_use> block must contain valid JSON with a "tool" field. Begin. Review the Knowledge Base Overview, identify the most relevant areas for the question, and start exploring. Knowledge Base Overview: { knowledge_base_overview } |