YouRA: A Persistent-State Architecture for Evidence-Traceable Autonomous Research Agents
Organizations: Electronics and Telecommunications Research Institute, Republic of Korea · Department of Artificial Intelligence, University of Science and Technology, Republic of Korea
Abstract
End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not maintained as persistent, verifiable state across long-horizon pipelines. We present YouRA (Your Research Agent), an architecture for stateful, evidence-traceable autonomous research. YouRA preserves research state, execution evidence, and failure history across the research trajectory by integrating three components: a Verification State Architecture (VSA) that tracks hypotheses, gates, and evidence pointers; an Independent Controller that turns state and reflection records into lifecycle, recovery, and debate/review control while separating control from execution; and Stateful Reflection that logs failures as structured lessons and routes recovery through bounded repair, redesign, or reset. On MLR-Bench's predefined ten-task end-to-end subset, YouRA improves over both MLR-Agent and AI Scientist V2 on scalar Overall across all three matched backbones. An automated diagnostic using MLR-Bench's hallucination taxonomy reports intersection/union counts for four fact-based failure types, and data-provenance diagnostic shows more real-data-based outputs. Ablating each of the four components (the VSA, the Independent Controller, MCP tool access, and reflection-guided recovery) supports their separable contributions. Removing either core-state component drops YouRA below the full system. Code: https://github.com/PrayPrey/Your-Research-Agent.
Figures & tables
| System | Backbone LLM | Clarity | Novelty | Soundness | Significance | Overall |
| MLR-Agent | Sonnet 4.5 | |||||
| Opus 4.5 | ||||||
| Sonnet 4.6 | ||||||
| AI Scientist V2 | Sonnet 4.5 | |||||
| Opus 4.5 | ||||||
| Sonnet 4.6 |
| Comparison | Backbone | Win | Tie | Lose |
| YouRA vs MLR-Agent | Sonnet 4.5 | 25 | 14 | 1 |
| Opus 4.5 | 26 | 5 | 9 | |
| Sonnet 4.6 | 21 | 14 | 5 | |
| YouRA vs AI Scientist V2 | Sonnet 4.5 | 18 | 14 | 8 |
| Opus 4.5 | 17 | 12 | 11 | |
| Sonnet 4.6 | 20 | 12 | 8 |
| Configuration | Backbone | VSA | IC | MCP | Refl. | Clarity | Novelty | Soundness | Significance | Overall |
| YouRA (full) | Sonnet 4.5 | ✓ | ✓ | ✓ | ✓ | |||||
| Sonnet 4.6 | ✓ | ✓ | ✓ | ✓ | ||||||
| Opus 4.5 | ✓ | ✓ | ✓ | ✓ | ||||||
| w/o VSA | Sonnet 4.5 | ✗ | ✓ | ✓ | ✓ | |||||
| Sonnet 4.6 | ✗ | ✓ | ✓ | ✓ | ||||||
| Opus 4.5 | ✗ | ✓ | ✓ | ✓ |
| Flag precision by category | Flag precision by system | Auditor agreement | ||||||
| Count | Rate | Count | Rate | Count | Rate | |||
| Hallucinated Methodology | 92/102 | 90.2% | YouRA | 68/90 | 75.6% | YouRA | 26/30 | 86.7% |
| Faked Experimental Results | 64/85 | 75.3% | MLR-Agent | 70/90 | 77.8% | MLR-Agent | 21/30 | 70.0% |
| Mathematical Errors | 15/21 | 71.4% | AI Scientist V2 | 65/90 | 72.2% | AI Scientist V2 | 22/30 | 73.3% |
| Nonexistent Citations | 32/62 | 51.6% | ||||||
| Overall | 203/270 | 75.2% | , | Overall | 69/90 | 76.7% | ||
| Flags | Precision | Corrected | |
| YouRA | 420 | 75.6% | 317 |
| MLR-Agent | 552 | 77.8% | 429 |
| AI Scientist V2 | 413 | 72.2% | 298 |
| 95% CI | |||
| YouRA MLR-Agent | 0.003 | ||
| YouRA AI Scientist V2 | 0.571 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Tool | Role | Lifecycle coverage | Usage in YouRA |
| Archon | Sequential Memory | Full lifecycle | RAG-backed knowledge base; stores implementation patterns; manages task lifecycle (todo/doing/review/done) |
| Exa | Evidence Search | Grounding, debate, planning, validation | Web-based evidence collection via Search, Contents, and Research APIs |
| Semantic Scholar | Literature Search | Grounding, debate, evidence, writing, review | Academic Graph and Recommendations APIs for paper metadata and citations |
| Serena | Code Analysis / Memory | Scoping, debate, hypothesis loop, writing, review | Code-aware symbol navigation; optional in Experiment Design when base code exists, stronger in Implementation Planning–Coding & Validation; memory support where explicitly invoked |
| Clear Thought | Structured Reasoning | Debate and evidence checking | Scientific method and mental models for hypothesis refinement |
| Workflow stage | Tools | Input | Output |
| Problem Scoping | Archon, Serena | Researcher’s interests | Concrete research question |
| Literature Grounding | Archon, SS, Exa | Research question | 3 research gaps |
| Hypothesis Debate | Archon, SS, Exa, Serena, ClearThought | 3 research gaps | Formalized hypothesis |
| Verification Planning | Archon, Exa, ClearThought | Formalized hypothesis | 3–7 sub-hypothesis DAG |
| Experiment Design | Archon, Exa, Serena | Sub-hypotheses | Experiment specification |
| Implementation Planning | Archon, Serena | Experiment spec | PRD, architecture, logic, config, tasks |
| Level | Target | Criterion |
| Repair | Coding & Validation (in-loop) | Local runtime, test, or fabricated-output failure repairable before archive-producing escalation |
| Redesign | Hypothesis Debate | MUST_WORK failure requiring hypothesis or mechanism revision after repair |
| Reset | Problem Scoping | Evidence contradiction indicating that the current hypothesis is unsupported |
| Backbone | Tasks | Total | Per task | Median | Range |
| Sonnet 4.5 | 10 | 28 | 2.8 | 2.5 | 0–8 |
| Opus 4.5 | 10 | 40 | 4.0 | 3.0 | 0–11 |
| Sonnet 4.6 | 10 | 18 | 1.8 | 1.5 | 0–4 |
| Task(MLR-Bench) | Sonnet 4.5 | Opus 4.5 | Sonnet 4.6 |
| bi_align | 2 | 3 | 2 |
| buildingtrust | 0 | 7 | 1 |
| data_problems | 3 | 3 | 0 |
| dl4c | 8 | 4 | 1 |
| mldpr | 3 | 6 | 1 |
| question | 7 | 11 | 4 |
| Backbone | redesign | reset | unclass. | Total |
| Sonnet 4.5 | 3 | 25 | 0 | 28 |
| Opus 4.5 | 1 | 39 | 0 | 40 |
| Sonnet 4.6 | 1 | 17 | 0 | 18 |
| Total | 5 | 81 | 0 | 86 |
| Comparison | Overall | Perm. | Holm | W/L/T | Sign | Backbone deltas |
| YouRA MLR-Agent | 0.006 | 0.012 | 18/12/0 | 0.362 | ||
| YouRA AI Scientist V2 | 0.044 | 0.044 | 20/8/2 | 0.036 |
| Stage | System | Backbone LLM | Consistency | Clarity | Novelty | Soundness | Feasibility | Significance | Overall |
| Idea | YouRA | Sonnet 4.5 | – | ||||||
| Opus 4.5 | – | ||||||||
| MLR-Agent | Sonnet 4.5 | – | |||||||
| Opus 4.5 | – | ||||||||
| Proposal | YouRA | Sonnet 4.5 | |||||||
| Opus 4.5 |
| Stage | Backbone | Overall | Perm. | Holm | YouRA higher | MLR-Agent higher |
| Idea | Sonnet 4.5 | 0.001 | 0.004 | – | Cons., Clar., Sig., Overall | |
| Idea | Opus 4.5 | 0.246 | 0.739 | Nov., Feas. | Sig. | |
| Proposal | Sonnet 4.5 | 0.020 | 0.099 | – | Clar., Sound. | |
| Proposal | Opus 4.5 | 0.428 | 0.857 | Cons., Nov., Feas. | Sound. |
| System | Backbone | Both analyzers | Opus 4.6 | GPT-5.4 |
| YouRA | Sonnet 4.5 | 8/10 | 8/10 | 9/10 |
| Opus 4.5 | 9/10 | 9/10 | 9/10 | |
| Sonnet 4.6 | 10/10 | 10/10 | 10/10 | |
| AI Scientist V2 | Sonnet 4.5 | 6/10 | 6/10 | 6/10 |
| Opus 4.5 | 8/10 | 9/10 | 8/10 | |
| Sonnet 4.6 | 3/10 | 3/10 | 4/10 |
| Removed component | Overall | Perm. | Holm | W/L/T | Sign | Backbone drops |
| MCP | 20/5/5 | 0.004 | ||||
| Reflection | 0.007 | 0.007 | 17/11/2 | 0.345 | ||
| VSA | 0.002 | 0.004 | 22/8/0 | 0.016 | ||
| Independent Controller | 0.001 | 0.003 | 22/6/2 | 0.004 |
| Drop-size contrast | drop | Perm. | Holm | W/L/T |
| MCP Reflection | 0.169 | 1.000 | 17/9/4 | |
| MCP VSA | 1.000 | 1.000 | 12/16/2 | |
| MCP Independent Controller | 1.000 | 1.000 | 12/15/3 | |
| Reflection VSA | 0.285 | 1.000 | 11/18/1 | |
| Reflection Independent Controller | 0.252 | 1.000 | 11/16/3 | |
| VSA Independent Controller | 1.000 | 1.000 | 17/9/4 |