RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation
Abstract
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success. We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction. RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation. After each training stage, it performs outcome-grounded backward attribution on low-outcome trajectories, aggregates recurrent and policy-controlled capability bottlenecks, and dynamically selects the corresponding process rewards for the next stage. Recurrent failures not covered by the existing capability space trigger a separate, controlled expansion procedure. We evaluate REWARDWEAVER on SOTOPIA, Amazon?HistoryPrice, and a newly constructed Sales Benchmark. Across social interaction, bilateral bargaining, and domain-specific sales, REWARDWEAVER establishes new state-of-the-art (SOTA) results. Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.
Figures & tables
| Self-Play | Partner | |||
| Method | SOTOPIA Goal / Overall | SOTOPIA-Hard Goal / Overall | SOTOPIA Goal / Overall | SOTOPIA-Hard Goal / Overall |
| Proprietary LLMs | ||||
| GPT-4o | 8.19 / 3.76 | 6.97 / 3.46 | 8.19 / 3.76 | 6.97 / 3.46 |
| Claude-3.5-Sonnet | 8.29 / 3.71 | 6.33 / 3.09 | 8.42 / 3.77 | 6.64 / 3.30 |
| DeepSeek-V3 | 8.15 / 3.62 | 6.34 / 3.09 | 8.14 / 3.72 | 6.69 / 3.31 |
| Large Reasoning Models | ||||
| Self-Play | Partner | |||
| Method | SOTOPIA Goal / Overall | SOTOPIA-Hard Goal / Overall | SOTOPIA Goal / Overall | SOTOPIA-Hard Goal / Overall |
| RewardWeaver (ours) | 8.72 / 4.47 | 7.90 / 4.38 | 8.49 / 4.47 | 7.66 / 4.22 |
| Dynamic Reward Adaptation | ||||
| Outcome-only GRPO | 8.43 / 4.31 | 7.53 / 4.05 | 8.35 / 4.26 | 7.71 / 4.12 |
| Static Top- Dimensions | 8.44 / 4.31 | 7.77 / 4.14 | 8.36 / 4.23 | 7.49 / 3.83 |
| Random Active Set | 8.47 / 4.41 | 7.56 / 4.18 | 8.33 / 4.27 | 7.31 / 3.97 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | # Rubrics | Behavioral Scope |
|---|---|---|
| Goal Expression and Persistence | 5 | Goal articulation and retention |
| Information Acquisition | 8 | Information elicitation and clarification |
| Information Use | 5 | Evidence use and adaptation |
| Concern Handling and Persuasion | 6 | Objection handling and negotiation |
| Goal Progression and Closure | 7 | Progress, strategy, and completion |
| Relationship Management | 5 | Empathy and interaction quality |
| ID | Capability | Definition | Positive Evidence | Negative Evidence |
|---|---|---|---|---|
| explicit-goal | Explicit Goal Expression | Explicitly state a task-required request, position, information, feeling, boundary, or correction that has not yet been expressed. | Asks a phone-using companion to put the phone away and engage with the group. | Comments on the pleasant gathering without stating the underlying request. |
| elicit-key-info | Targeted Information Elicitation | Directly ask for task-relevant information that is required but not yet available. | Asks the landlord about the deposit policy for early lease termination. | Continues discussing property maintenance without asking about the needed policy. |
| use-key-information | Use of Key Information | Accurately use at least one known, relevant, and safely usable fact to support the current response. | Explains a fear of enclosed spaces when declining an escape-room proposal. | Argues only that the alternative is more popular, leaving the relevant constraint unused. |
| adapt-to-new-info | Adapt to New Information | Adjust a request or plan when newly disclosed information changes its feasibility. | Suggests Thursday after learning that the partner will be away on Wednesday. | Repeats the Wednesday request despite the disclosed trip. |
| address-concern | Address Concerns with New Options | After the partner raises a specific concern, propose an actionable adjustment that directly reduces that concern. | Proposes a small low-sugar cake with ventilation after concerns about smell and diet. | Repeats that baking is enjoyable without modifying the proposal. |
| strategy-pivot | Switch Strategies | After repeated rejection or deferral, use a new argument, angle, or plan rather than repeating the same request. | Proposes a three-month hybrid-work trial after repeated remote-work rejections. | Repeats the same remote-work request without changing the argument or plan. |
| Field | Operationalization |
|---|---|
| Rubric ID | fixed-use-key-information-v1 |
| Applicability | The Rubric applies when at least one safe, directly relevant fact, preference, constraint, or confirmed arrangement is available in the visible context before the target response. |
| Score | The target response accurately uses at least one currently relevant item of known information to explain, request, confirm, adapt, or respond. |
| Score | The Rubric applies, but the target response uses none of the relevant known information, or misstates a relevant fact, quantity, person, relation, or constraint. |
| Non-applicable ( ) | No safe and directly relevant information is available for the current response. Failure to use available information is scored as , not . |
| Abstain ( ) | The visible evidence is insufficient to determine applicability or scoring without guessing. |
| Family | Scenario |
|---|---|
| R01 | Staged corporate procurement |
| R02 | Data-use purpose separation |
| R03 | Cross-city service-network due diligence |
| R04 | Multi-site acceptance and objection terms |
| R05 | Corporate lead qualification |
| R06 | Sign-language and bilingual support |
| Method | MI Deal Rate | CI Correct No-Deal Rate | Task Success Rate | Normalized Seller Utility | Expected MI Seller Return |
|---|---|---|---|---|---|
| Proprietary LLMs | |||||
| doubao-seed-2-0-pro | 94.40 | 100.00 | 94.50 | 0.85 | 0.80 |
| glm-5.3-flash | 95.20 | 100.00 | 95.30 | 0.91 | 0.86 |
| glm-5.2 | 92.00 | 100.00 | 92.20 | 0.93 | 0.86 |
| qwen3.8-max | 80.80 | 100.00 | 81.30 | 0.95 | 0.76 |
| qwen3.8-flash | 83.20 | 100.00 | 83.60 | 0.93 | 0.77 |
| Self-Play | Partner | |||
| Method | Sales Goal / Overall | Sales-Hard Goal / Overall | Sales Goal / Overall | Sales-Hard Goal / Overall |
| Proprietary LLMs | ||||
| doubao-seed-2-1-pro | 8.55 / 4.36 | 7.29 / 3.99 | 8.81 / 4.43 | 7.47 / 4.10 |
| qwen3.8-max | 8.65 / 4.38 | 7.76 / 4.09 | 8.76 / 4.43 | 7.76 / 4.17 |
| gemini-3.1-pro-preview | 8.63 / 4.37 | 7.29 / 4.11 | 8.83 / 4.48 | 7.59 / 4.08 |
| deepseek-v4-pro | 8.68 / 4.30 | 7.24 / 3.97 | 8.67 / 4.39 | 7.82 / 4.17 |
| Analysis | Spearman |
| Bottleneck importance vs. | 0.28 |
| Current failure rate vs. | 0.21 |
| vs. , controlling for | 0.22 |
| Within-stage association | 0.31 |
| Self-Play | Partner | |||
|---|---|---|---|---|
| SOTOPIA | SOTOPIA-Hard | SOTOPIA | SOTOPIA-Hard | |
| 2 | 8.46 / 4.32 | 7.50 / 3.97 | 8.38 / 4.17 | 7.46 / 3.92 |
| 4 | 8.31 / 4.29 | 7.79 / 4.16 | 8.21 / 4.19 | 7.37 / 4.03 |
| 6 | 8.72 / 4.47 | 7.90 / 4.38 | 8.49 / 4.47 | 7.66 / 4.22 |
| 8 | 8.49 / 4.44 | 7.67 / 4.25 | 8.31 / 4.25 | 7.59 / 4.15 |
| 10 | 8.50 / 4.46 | 7.63 / 4.36 | 8.34 / 4.37 | 7.57 / 4.21 |
| Self-Play | Partner | |||
| Interval | SOTOPIA | SOTOPIA-Hard | SOTOPIA | SOTOPIA-Hard |
| 10 | 8.28 / 4.41 | 7.73 / 4.36 | 8.32 / 4.33 | 7.59 / 4.20 |
| 20 | 8.50 / 4.47 | 7.81 / 4.39 | 8.27 / 4.19 | 7.47 / 4.03 |
| 30 | 8.72 / 4.47 | 7.90 / 4.38 | 8.49 / 4.47 | 7.66 / 4.22 |
| 60 | 8.21 / 4.26 | 7.44 / 3.95 | 8.34 / 4.28 | 7.49 / 4.04 |
| Capability | Stage 1 | Stage 2 | Stage 3 | Stage 4 | Stage 5 | Stage 6 | Stage 7 |
| Address Concerns with New Options | 0.80 | 0.85 | 0.92 | 0.95 | 0.86 | 0.89 | 0.90 |
| State Benefits | 0.84 | 0.88 | 0.97 | 0.98 | 0.95 | 0.94 | 0.98 |
| Conditional Concessions | 0.85 | 0.90 | 0.98 | 0.93 | 0.86 | 0.90 | 0.96 |
| Maintain Goal Commitment | 0.96 | 0.95 | 1.00 | 0.98 | 0.96 | 1.00 | 1.00 |
| Regulate Pressure | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| Switch Strategies | 0.62 | 0.91 | 1.00 | 1.00 | 0.88 | 1.00 | 1.00 |