AdaptEvo: Adaptive Agent Learning with Evolving Supervision
Organizations: Fudan University · Sun Yat-sen University · Xiaohongshu
Abstract
Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence. Written rules can leave gaps in decision guidance and process evaluation, while reference judgments vary in their support from the rules and evidence. To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics. Its Training module uses Confidence-Adaptive GRPO (CA-GRPO) to balance outcome and process rewards according to reference confidence. Its Evolution module synthesizes reusable decision knowledge from recurring failures across training cases and refines process rubrics to detect overlooked errors. To support empirical evaluation, we construct an industrial multimodal content moderation dataset comprising a training set and In-Period and Out-of-Period test sets, with the latter collected under changed rules. Using Qwen3.6-35B-A3B, AdaptEvo achieves 61.9% exact-label accuracy and 72.2% binary decision accuracy on In-Period, exceeding GRPO by 7.5 and 3.7 percentage points, respectively. On Out-of-Period, the policy trained with CA-GRPO retains exact-label accuracy gains over the base model across evaluated checkpoints without injected decision knowledge, while GRPO declines with continued training. CA-GRPO also outperforms the tested fixed reward mixtures on both Out-of-Period metrics.
Figures & tables
| Model / Method | Knowledge | In-Period | Out-of-Period | ||
|---|---|---|---|---|---|
| ELA | BDA | ELA | BDA | ||
| Qwen3.5-4B | None | 46.8 | 60.8 | 50.8 | 64.5 |
| Qwen3.5-9B | None | 46.2 | 63.2 | 50.4 | 65.2 |
| Qwen3.5-27B | None | 55.1 | 68.5 | 53.8 | 66.0 |
| Qwen3.8-Flash-Next | None | 54.9 | 69.1 | 54.4 | 67.6 |
| GLM5.3-Flash | None | 56.3 | 69.6 | 55.6 | 67.5 |
| In-Period | Out-of-Period | |||
|---|---|---|---|---|
| Reward Mixing | ELA | BDA | ELA | BDA |
| Fixed (0.8:0.2) | 54.3 | 67.6 | 53.6 | 66.1 |
| Fixed (0.5:0.5) | 54.2 | 67.9 | 54.0 | 66.2 |
| Fixed (0.2:0.8) | 55.6 | 67.3 | 54.8 | 65.8 |
| CA-GRPO | 55.4 | 68.1 | 55.2 | 67.3 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Instance and correspondence |
|---|---|
| Post and queue | A mint-green eyeshadow post with product/swatch tags, nine post images, and a transaction-linked commercial marker for Brand A. Author metadata also supplies an avatar and background image (11 image inputs in total). The assigned queue reviews unfair competition in posts. |
| Candidates | , where , , , and . Here abbreviate this queue’s violation labels. |
| Queue rule | Establish commercial context using account information and commercial markers; assess the post and relevant author-originated context under the applicable label rules. This queue-wide guidance is supplied in full at initialization. |
| Label rules | For , commercial applicability must be accompanied by covered unfair comparison, such as unsupported scoring/ranking or one-sided competitor criticism; balanced multi-product reviews can be exempt. For , assess competitive denigration and its prerequisites and exemptions. For , assess consumer negative claims, with exemptions for specified subjective preferences or individual fit. Only previews are initially shown; the full rules are retrieved below. |
| Rule lookup | get_detail_rule(labels=[ ]) returns the full rules for all three violation labels in the current queue. The actual arguments are the corresponding label strings. |
| Evidence lookup | get_commercial_detail({}) returns no associated product items. The initial transaction-linked marker remains part of the evidence; this empty result does not establish that the post is noncommercial. |
| Agr. | S | A | U | % | |
|---|---|---|---|---|---|
| 4-0 | 1,968 | 473 | 245 | 2,686 | 79.12 |
| 3-1 | 287 | 175 | 212 | 674 | 19.85 |
| 2-2 | 11 | 7 | 6 | 24 | 0.71 |
| 2-1-1 | 5 | 4 | 1 | 10 | 0.29 |
| 1-1-1-1 | 0 | 0 | 1 | 1 | 0.03 |
| Total | 2,271 | 659 | 465 | 3,395 | 100.00 |
| Setting | Training pool | Supervision and evolution |
|---|---|---|
| Base Model | — | No task-specific RL |
| GRPO | Full pool | Reference-based outcome rewards |
| GRPO (High-Conf.) | High-confidence subset | Reference-based outcome rewards |
| CA-GRPO | Full pool | Confidence-adaptive outcome/process mixture; no injected decision knowledge; fixed rubric |
| AdaptEvo | Full pool | Continued CA-GRPO after joint knowledge/rubric evolution |
| Method | Knowledge | In-Period | Out-of-Period | ||
|---|---|---|---|---|---|
| ELA | BDA | ELA | BDA | ||
| Base model | None | 53.9 | 66.1 | 53.3 | 65.9 |
| GRPO | None | 54.4 | 68.5 | 51.9 | 65.2 |
| GRPO (High-Conf.) | None | 53.3 | 68.8 | 52.9 | 65.5 |
| Fixed (0.8:0.2) | None | 54.3 | 67.6 | 53.6 | 66.1 |
| Fixed (0.5:0.5) | None | 54.2 | 67.9 | 54.0 | 66.2 |
| GRPO | GRPO (High-Conf.) | CA-GRPO | ||||
| Training steps | ELA | BDA | ELA | BDA | ELA | BDA |
| Base | 53.3 | 65.9 | 53.3 | 65.9 | 53.3 | 65.9 |
| 10 | 54.5 | 66.0 | 55.1 | 66.8 | 54.1 | 66.2 |
| 20 | 54.2 | 66.2 | 54.4 | 67.2 | 54.3 | 65.7 |
| 30 | 54.2 | 66.5 | 53.9 | 66.8 | 54.5 | 65.9 |
| 40 | 53.3 | 65.9 | 52.0 | 65.1 | 54.4 | 65.7 |
| In-Period | In-Period + knowledge | Out-of-Period | ||||
|---|---|---|---|---|---|---|
| Outcome : process | ELA | BDA | ELA | BDA | ELA | BDA |
| 0.8:0.2 | 54.3 | 67.6 | 57.7 | 68.9 | 53.6 | 66.1 |
| 0.5:0.5 | 54.2 | 67.9 | 58.1 | 70.1 | 54.0 | 66.2 |
| 0.2:0.8 | 55.6 | 67.3 | 58.6 | 69.3 | 54.8 | 65.8 |
| [0pt][0pt] CA-GRPO (adaptive) | 55.4 | 68.1 | 59.1 | 71.3 | 55.2 | 67.3 |
| Internal alias | Main-table name | Knowledge |
|---|---|---|
| base | Qwen3.6-35B-A3B (base) | None |
| gtdirect-full | GRPO | None |
| gtdirect-support | GRPO (High-Conf.) | None |
| gtconf-v1 , library off | CA-GRPO | None |
| gtconf-v1 , library on | CA-GRPO | Evolved |
| gtconf-v2 | AdaptEvo | Evolved |