RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent
Organizations: Huawei Technologies, Co., Ltd.
Abstract
Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels. We propose RefCon, which combines sequential self-refinement with parallel self-contrast to extract higher-quality memories without gold labels. Evaluated on AppWorld and BFCL-V3 across multiple context-evolving agent frameworks, RefCon delivers strong and consistent gains, including relative improvements of 21.6% on ACE and 16.6% on ReMe over no-scaling baselines, while a diversity-focused variant (DivCon) achieves a 35.5% gain on ReasoningBank. RefCon consistently outperforms existing baselines without ground-truth labels, and generalizes across model scales and to software engineering tasks, where it surpasses even ground-truth baselines. We further analyze the accuracy-token trade-off and scaling behavior, showing RefCon maintains favorable efficiency and continues to improve as more trajectories are used, unlike diversity-only scaling which saturates earlier.
Figures & tables
| Methods | AppWorld | BFCL-V3 | ||||
| Avg@3 | Pass@3 | Avg@3 | Pass@3 | |||
| TGC | SGC | TGC | SGC | |||
| Baseline | ||||||
| ReAct | 87.72 | 68.42 | 62.00 | |||
| With Gold Labels/Ground-Truth | ||||||
| ACE (with Ground-Truth) | 96.49 | 89.47 | 82.00 | |||
| Methods | AppWorld | |||
| Avg@3 | Pass@3 | |||
| TGC | SGC | TGC | SGC | |
| ACE | ||||
| ACE (Without Scaling) | 77.19 | 73.68 | ||
| ACE (Parallel Scaling) | 84.21 | 78.95 | ||
| ACE (Sequential Scaling) | 84.21 | 78.95 | ||
| Method | Iter 1 | Iter 2 | Iter 3 |
|---|---|---|---|
| Methods Without Refinement | |||
| ReAct | — | — | |
| RB (Par) | — | — | |
| ACE (w/ GT) | — | — | |
| Methods With Refinement | |||
| Self-Refine | |||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Methods | AppWorld | BFCL-V3 | ||||
| Avg@3 | Pass@3 | Avg@3 | Pass@3 | |||
| TGC | SGC | TGC | SGC | |||
| Baseline | ||||||
| ReAct | 92.98 | 78.95 | 72.00 | |||
| With Gold Labels/Ground-Truth | ||||||
| ACE (with Ground-Truth) | 96.49 | 89.47 | 84.00 | |||
| Methods | AppWorld | BFCL-V3 | ||||
| Avg@3 | Pass@3 | Avg@3 | Pass@3 | |||
| TGC | SGC | TGC | SGC | |||
| Baseline | ||||||
| ReAct | 29.82 | 15.79 | 72.00 | |||
| With Gold Labels/Ground-Truth | ||||||
| ACE (with Ground-Truth) | 29.82 | 15.79 | 80.00 | |||
| Methods | AppWorld | |||
| Avg@3 | Pass@3 | |||
| TGC | SGC | TGC | SGC | |
| ACE | ||||
| First Run | 85.96 | 68.42 | ||
| Second Run (First Refinement) | 92.98 | 84.21 | ||
| Third Run (Second Refinement) | 92.98 | 78.95 | ||
| Methods | AppWorld | |||
| Avg@3 | Pass@3 | |||
| TGC | SGC | TGC | SGC | |
| ACE | ||||
| First Run | 91.23 | 78.95 | ||
| Second Run (First Alternative) | 89.47 | 73.68 | ||
| ReasoningBank | ||||
| Section: strategies_and_hard_rules |
| Content: For Spotify authentication, always call apis.supervisor.show_account_passwords() to retrieve stored credentials instead of creating dummy passwords. |
| Section: strategies_and_hard_rules |
| Content: When collecting songs across multiple libraries (songs, albums, playlists), use set() to deduplicate song_ids from all sources before processing. |
| Section: strategies_and_hard_rules |
| Title: Authenticate Before Accessing Private Data |
| Description: Always authenticate with the service before attempting to access private user data like libraries. |
| Content: When working with music streaming services, private library data (songs, albums, playlists) requires authentication. First explore available APIs, then use stored credentials to obtain an access token before making any data retrieval requests. |
| Title: Handle Pagination for Complete Data Retrieval |
| Description: Implement pagination loops to retrieve all items from paginated API responses. |
| Content: Many APIs return data in pages with limits. Use while loops that increment page_index until an empty response is returned, collecting all items across pages to ensure complete data retrieval. |
| Original Memory |
|---|
| When To Use: When working with paginated API responses |
| Content: All trajectories successfully implemented pagination with while loops and page_index incrementing. This consistent pattern across all successful trajectories shows that proper pagination handling is critical for complete data retrieval. The pattern of checking for empty responses to break the loop is reliable. |
| When To Use: When handling authentication tokens in multi-step workflows |
| Content: All trajectories successfully extracted and reused the access_token from login responses across multiple API calls. The consistent pattern of storing spotify_access_token = login_result[’access_token’] and passing it to subsequent calls demonstrates proper token management. None of the trajectories encountered token expiration issues, suggesting the token lifespan was sufficient for the workflow duration. |
| Retrieved Memory (Sequential Scaling) |
|---|
| When To Use: When you encounter verification issue in Venmo |
| Content: If you encounter issues with optional verification steps (like checking Venmo status), proceed with the main task (sending messages) to ensure completion. The primary goal of delivering the message is more important than perfect filtering, especially when verification may be unreliable. |
| When To Use: When filtering contacts by a condition and finding no matches, but the user has requested to send a specific message |
| Content: If no recipients match the specified filter criteria, consider alternative interpretations of the user’s request before concluding the task. The message content itself may provide clues about the intended audience. |
| Retrieved Memory (RefCon) |
| Sequential Scaling (Failed) | RefCon (Successful) |
|---|---|
| Assistant: contacts_without_venmo = [] for contact in all_contacts: email = contact[’email’] search_result = apis.venmo.search_users( access_token=venmo_access_token, query=email, page_limit=1) if not search_result: contacts_without_venmo.append(contact) print(contacts_without_venmo) Environment: [] Assistant: contacts_without_venmo = [] for contact in all_contacts: email = contact[’email’] search_result = apis.venmo.search_users( access_token=venmo_access_token, query=email) if not search_result: contacts_without_venmo.append(contact) print(contacts_without_venmo) Environment: [] Assistant: contacts_without_venmo = [] for contact in all_contacts: full_name = f"{contact[’first_name’]} {contact[’last_name’]}" search_result = apis.venmo.search_users( access_token=venmo_access_token, query=full_name) if not search_result: contacts_without_venmo.append(contact) print(contacts_without_venmo) Environment: [] | Assistant: all_contacts = siblings + roommates contacts_without_venmo = [] for contact in all_contacts: search_result = apis.venmo.search_users( access_token=venmo_access_token, query=contact[’email’]) if not search_result: contacts_without_venmo.append(contact) print(contacts_without_venmo) Environment: [] Assistant: print(apis.venmo.search_users( access_token=venmo_access_token, query=’kevin.blac@gmail.com’)) Environment: [{"first_name": "Kevin", "email": "kevin.blac@gmail.com", …}] Assistant: all_contacts = siblings + roommates contacts_without_venmo = [] for contact in all_contacts: search_result = apis.venmo.search_users( access_token=venmo_access_token, query=contact[’email’]) found = any(user[’email’] == contact[’email’] for user in search_result) if not found: contacts_without_venmo.append(contact) print(contacts_without_venmo) Environment: [{"contact_id": 801, "first_name": "Martin", …}] |
| Step | Successful Trajectory | Failed Trajectory |
|---|---|---|
| 1–2 | (Identical) Retrieve account passwords; log in to Spotify to obtain access_token . | |
| 3–7 | (Identical) Inspect API schemas for show_song , show_song_privates , show_song_library , show_album_library , show_playlist_library . | |
| 8–11 | (Identical) Collect 79 unique song IDs from all three libraries; filter to 18 R&B songs via show_song() genre check. | |
| 12 | Read play_count from show_song ⬇ rb_songs_with_play_count = [] for song_id in all_song_ids : song_info = apis . spotify . show_song ( song_id = song_id ) if song_info and ’ r & b ’ in song_info [’ genre ’]. lower (): rb_songs_with_play_count . append ({ ’ title ’: song_info [’ title ’], ’ play_count ’: song_info [’ play_count ’] }) show_song returns a rich public schema including play_count : {song_id, title, album_id, duration, artists, genre, play_count , rating, ...} Genre filtering and play count retrieval are done in a single loop . All 18 R&B songs receive their correct play counts. | Read play_count from show_song_privates ⬇ rb_songs_with_play_count = [] for song in rb_song_details : song_id = song [’ song_id ’] private_info = apis . spotify . show_song_privates ( song_id = song_id , access_token = spotify_access_token ) if private_info : play_count = private_info . get (’ play_count ’, 0) rb_songs_with_play_count . append ({ ’ title ’: song [’ title ’], ’ play_count ’: play_count }) show_song_privates only exposes per-user interaction flags: {liked, reviewed, in_song_library, downloaded} There is no play_count field. The .get(’play_count’, 0) fallback silently returns 0 for all 18 songs , producing no error and no warning. |
| 13 | Sorted by real play counts: ⬇ [" Mysteries of the Silent Sea ", " Crimson Veil ", " Haunted Memories ", " Fire and Ice "] | All counts equal 0; order is arbitrary: ⬇ [" Shadows of the Past ", " When Fate Becomes a Foe ", " The Curse of Loving You ", " Lost in a Moment ’ s Grace "] |
| Self-Contrast Prompt |
|---|
| You are an expert AI analyst comparing multiple step sequences which might be successful or failed to extract differential insights. |
| Your task is to compare and contrast these trajectories to identify the most useful and generalizable strategies as memory items using self-contrast reasoning. |
| Focus on critical decision points, technique variations, and approach differences. |
| COMPARATIVE ANALYSIS FRAMEWORK: |
| - DECISION CONTRAST: Compare critical decisions made in success vs failure cases |
| - TECHNIQUE VARIATIONS: Identify different approaches and their outcomes |
| Self-Refine Prompt for AppWorld |
|---|
| Let’s carefully re-examine the previous trajectory, including your reasoning steps and action taken. Pay special attention to whether you used the best API sequence and whether you used the API correctly. If you find inconsistencies, correct them. If everything seems correct, make it more efficient. Now, solve the same problem again from scratch. |
| Self-Diversity Prompt for AppWorld |
| Below is the previous trajectory, the solution might be correct or wrong. Now solve the same problem using a DIFFERENT reasoning approach. Focus on exploring alternative strategies. |
| Critique Trajectory Prompt for Self-Refine |
|---|
| You are an expert reviewer analyzing an AI assistant’s multi-turn tool-calling trajectory. Your job is to identify mistakes, missed actions, and suboptimal decisions. For each turn in the trajectory, evaluate: |
| 1. Did the assistant call the appropriate tools? If a user requested an action (e.g., book, cancel, update), did the assistant actually make a tool call, or did it just respond with text? |
| 2. Were the tool arguments correct? Check for wrong parameter values, missing required arguments, or arguments that contradict the user’s request. |
| 3. Did the assistant use information from previous tool responses correctly? For example, if a lookup returned an ID, did the assistant use that ID in subsequent calls? |
| 4. Were there any unnecessary or redundant tool calls? |
| 5. Did the assistant follow the logical sequence of operations? (e.g., lookup before booking, authenticate before accessing protected resources) |
| Critique Trajectory Prompt for Self-Refine |
|---|
| You are an expert reviewer analyzing an AI assistant’s multi-turn software engineering trajectory. Your job is to identify mistakes, missed files, and suboptimal debugging decisions. For each turn in the trajectory, evaluate: |
| 1. Did the assistant explore the repository effectively? Did it locate the relevant source files and classes, or did it waste turns on unrelated directories? |
| 2. Was the bug localization accurate? Check if the assistant correctly identified the root cause of the issue before attempting a fix. |
| 3. Did the assistant use the environment and test tools correctly? For example, if a reproduction script was created, did the assistant analyze the output to guide the patch? |
| 4. Was the generated patch functional and minimal? Identify if the assistant introduced unnecessary changes or failed to follow the repository’s coding style. |
| 5. Did the assistant follow a logical debugging sequence? (e.g., search reproduce fix verify via tests) |