cs.AISep 27, 2026

Self-Designed Evaluators and Warm Memory for Long-Horizon Agents

Authors: Saeid Asgari, Emre Kiciman, Leonardo de Oliveira Nunes, Ranveer Chandra

Organizations: Microsoft

Abstract

A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

    May 12, 2026Di Wu, Zixiang Ji, Asmi Kawatkar +4Long-Term Agent MemoryLong-Term Memory

  2. EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

    Jul 6, 2026Xingze Gao, Chuanrui Hu, Hongda Chen +9Self-Evolving AgentsAgentic Benchmarks

  3. Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

    Aug 13, 2026Yiwei Li, Wanli Yang, Hexiang Tan +10Long-Horizon AgentsAutonomous Agents