Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
Organizations: Gaoling School of Artificial Intelligence Renmin University of China Beijing, China
Abstract
As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Preference | Family | Trigger | Requirement |
|---|---|---|---|
| Private note format | In-support | Venmo payment | The note follows a private bracketed format; its code is determined by the payment amount. |
| SMS sign-off | In-support | Text message | The message ends with the user’s first name. |
| SMS character checksum | Out-of-support | Text message | The message ends with a checksum equal to its character count modulo seven. |
| Habitual card | Out-of-support | Venmo payment | The first attempted payment card is the user’s latent habitual bank card. |
| Running spend total | Out-of-support | Venmo payment | The note ends with the cumulative amount successfully sent to that recipient, including the current payment and earlier successful ones. |
| Preference | Trigger | Requirement |
|---|---|---|
| SMS sign-off | Text message | End with the user’s first name. |
| SMS greeting | Text message | Begin with a greeting word. |
| SMS terseness | Text message | Use at most five words. |
| Payment has note | Venmo payment | Include a non-empty description. |
| Payment-note lowercase | Venmo payment | Write the note entirely in lowercase. |
| Single-word payment note | Venmo payment | Use exactly one word. |
| Preference | Feedback role | Requirement |
|---|---|---|
| SMS greeting | Carried | Begin with a greeting word. |
| SMS sign-off | Carried | End a text message with the user’s first name. |
| Venmo private | Carried | Mark a Venmo transaction private. |
| Payment-note category | Target withheld | Begin the note with the literal category tag [personal] . |
| Payment-note initials | Target withheld | End the note with parenthesized initials such as (J.D.) . |