ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
Organizations: University of Hong Kong · The Chinese University of Hong Kong · Jiangxi Science and Technology Normal University · Shenzhen University
Abstract
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.
Figures & tables
| Stage | Actions |
|---|---|
| Query formulation | Original question; Query2Doc; two-query decomposition |
| Retriever | BM25; E5 dense; BGE-M3 hybrid; Qwen3 embedding; reciprocal-rank fusion |
| Passage selection | Relevance ranking; maximal marginal relevance |
| Retrieval width | 3 or 6 passages |
| Retrieval control | Stop; continue while fewer than three rounds have executed |
| Reranking | Preserve order; Qwen3 reranker |
| Method | Workflow utility | Answer F1 | Execution words |
|---|---|---|---|
| Base | 0.5648 | 0.6954 | 5,348.7 |
| PPO | 0.5651 0.0021 | 0.6913 0.0021 | 5,170.4 76.8 |
| VinePPO | 0.5939 0.0062 | 0.7115 0.0026 | 4,817.7 179.7 |
| ASCT-Uniform | 0.6146 0.0191 | 0.7316 0.0168 | 4,792.2 151.4 |
| ASCT-UCT | 0.6075 0.0241 | 0.7257 0.0260 | 4,841.2 112.8 |
| ASCT-AgentUCT | 0.6187 0.0076 | 0.7266 0.0126 | 4,420.8 213.3 |
| Token count (M) | VinePPO | ASCT- Uniform | ASCT- UCT | ASCT- AgentUCT |
|---|---|---|---|---|
| Auxiliary total | 388.374 4.694 | 197.077 0.024 | 197.588 0.068 | 193.088 0.404 |
| Auxiliary environment execution | ||||
| Environment subtotal | 137.088 1.231 | 197.077 0.024 | 197.588 0.068 | 193.088 0.404 |
| Generator input (4B) | 89.972 0.278 | 139.452 0.162 | 139.613 0.193 | 137.776 0.174 |
| Generator output (4B) | 3.590 0.009 | 5.874 0.003 | 5.883 0.002 | 5.833 0.003 |
| Query embedding (0.6B) | 0.611 0.012 | 0.689 0.002 | 0.690 0.001 | 0.689 0.002 |
| Evaluator | Mean rollout | Auxiliary tokens (M) | |
|---|---|---|---|
| VinePPO | 0.5507 0.0027 | 388.374 4.694 | 0.4893 0.0019 |
| ASCT-Uniform | 0.5283 0.0022 | 197.077 0.024 | 0.4972 0.0022 |
| ASCT-UCT | 0.5515 0.0031 | 197.588 0.068 | 0.5203 0.0031 |
| ASCT-AgentUCT | 0.5526 0.0016 | 193.088 0.404 | 0.5220 0.0016 |
| Dataset | Method | F1 | Utility | Execution words |
|---|---|---|---|---|
| 2Wiki | Base | 0.5583 | 0.4732 | 3,488.0 |
| 2Wiki | PPO | 0.5883 0.0100 | 0.5067 0.0102 | 3,342.6 130.9 |
| 2Wiki | VinePPO | 0.6517 0.0375 | 0.5767 0.0360 | 3,069.3 133.7 |
| 2Wiki | ASCT-Uniform | 0.6567 0.0321 | 0.5832 0.0329 | 3,009.9 47.6 |
| 2Wiki | ASCT-UCT | 0.6583 0.0522 | 0.5837 0.0508 | 3,055.2 59.8 |
| 2Wiki | ASCT-AgentUCT | 0.6700 0.0321 | 0.5982 0.0313 | 2,939.3 73.7 |
| Planner | Type + correct relation | Type only | Name only | Type + opposite relation |
|---|---|---|---|---|
| Base | 0.820 | 0.010 | 0.000 | 0.020 |
| PPO | 0.830 0.044 | 0.033 0.042 | 0.000 0.000 | 0.020 0.000 |
| VinePPO | 0.947 0.035 | 0.280 0.108 | 0.000 0.000 | 0.040 0.010 |
| ASCT-Uniform | 0.960 0.017 | 0.300 0.111 | 0.000 0.000 | 0.037 0.006 |
| ASCT-UCT | 0.933 0.021 | 0.197 0.068 | 0.000 0.000 | 0.027 0.012 |
| ASCT-AgentUCT | 0.950 0.036 | 0.373 0.205 | 0.000 0.000 | 0.050 0.020 |
| Continuation variant | Test F1 | Test Utility | Aux. tokens (M) |
|---|---|---|---|
| ASCT-ActorRollout | |||
| ASCT-Uniform |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Inputs: actor , training tasks , evaluator , budget , utility . | |
|---|---|
| 1 | for each training iteration do |
| 2 | Freeze ; initialize actor-record buffer . |
| 3 | for each task in do |
| 4 | Initialize workflow state . |
| 5 | while is nonterminal do |
| 6 | Compute for ; sample . |
| Setting | Value |
|---|---|
| Planner | Qwen3-4B-Instruct-2507 |
| Adapter | LoRA rank 4, alpha 8, dropout 0; query and value projections |
| Optimizer | AdamW; learning rate |
| Optimizer batch size | 8 |
| PPO | Two update epochs per iteration; clipping |
| Advantage processing | Standardization across collected decisions within each iteration |
| Method | Credit at a sampled actor decision | Auxiliary evaluation |
|---|---|---|
| Base | No policy update | None |
| PPO | Terminal trajectory utility | None |
| VinePPO | Adjacent Monte Carlo state values | Actor continuations |
| ASCT-Uniform | Actor-centered action advantage | Uniform tree evaluator |
| ASCT-UCT | Actor-centered action advantage | Utility-directed UCT evaluator |
| ASCT-AgentUCT | Actor-centered action advantage | Cost-aware AgentUCT evaluator |
| Role / input | Core instruction | Output constraint |
|---|---|---|
| Actor / current workflow observation | “You select the next action in a RAG workflow.” | JSON with one action field; probabilities normalized over legal labels. |
| Query2Doc / question | “Write a short hypothetical passage that would answer the question.” | JSON with one pseudo_document string, used as the retrieval query. |
| Decomposition / question | “Decompose the multi-hop question into exactly two focused search queries.” | JSON queries : exactly two nonempty strings. |
| Answer / question and evidence passages | “Answer only from the evidence. Return the shortest answer span that satisfies the question.” | JSON answer and citations ; use UNKNOWN if evidence is insufficient and cite relevant evidence. |
| Evaluator | Total tokens (M) | Tokens / actor state | |
|---|---|---|---|
| ASCT-Uniform | 197.077 0.024 | 3,732.5 3.2 | 0.4972 0.0022 |
| ASCT-UCT | 197.588 0.068 | 3,749.5 2.6 | 0.5203 0.0031 |
| ASCT-AgentUCT | 193.088 0.404 | 3,664.5 9.2 | 0.5220 0.0016 |
| Method | |||
|---|---|---|---|
| PPO | 0.5651 0.0021 | 0.5651 0.0021 | 0.5651 0.0021 |
| VinePPO | 0.2055 0.0046 | 0.5162 0.0056 | 0.5550 0.0059 |
| ASCT-Uniform | 0.4175 0.0190 | 0.5752 0.0191 | 0.5949 0.0191 |
| ASCT-UCT | 0.4099 0.0241 | 0.5680 0.0241 | 0.5878 0.0241 |
| ASCT-AgentUCT | 0.4256 0.0075 | 0.5801 0.0076 | 0.5994 0.0076 |
| Method | Easy ( ) | Medium ( ) | Hard ( ) |
|---|---|---|---|
| Base | 0.6348 | 0.5723 | 0.4582 |
| PPO | 0.6405 0.0023 | 0.5702 0.0050 | 0.4616 0.0053 |
| VinePPO | 0.6613 0.0046 | 0.6031 0.0108 | 0.4836 0.0191 |
| ASCT-Uniform | 0.6565 0.0045 | 0.6334 0.0265 | 0.4968 0.0097 |
| ASCT-UCT | 0.6540 0.0056 | 0.6222 0.0345 | 0.5005 0.0097 |
| ASCT-AgentUCT | 0.6583 0.0111 | 0.6338 0.0178 | 0.5174 0.0120 |
| Seed | All decisions | Multi-action decisions | Sign changes | Rate (%) |
|---|---|---|---|---|
| 11 | 52,828 | 12,000 | 1,043 | 8.69 |
| 23 | 52,833 | 12,000 | 1,021 | 8.51 |
| 37 | 52,741 | 12,000 | 1,053 | 8.78 |
| Evaluator | Repeat Q SD | Credit agreement (%) | Tokens (k) |
|---|---|---|---|
| Root-stratified MC | |||
| ASCT-Uniform |
| Evaluator | Repeat Q SD | Credit (%) | Tokens (k) | |
|---|---|---|---|---|
| 12 | Root-stratified MC | |||
| 12 | ASCT-Uniform | |||
| 48 | Root-stratified MC | |||
| 48 | ASCT-Uniform |
| Evaluator | Repeat Q SD | Ref. order (%) | Tokens (k) |
|---|---|---|---|
| Root-stratified MC | |||
| ASCT-Uniform | |||
| ASCT-UCT | |||
| ASCT-AgentUCT | |||
| Actor MC |
| S/E/R | Tokens@12 (k) | Tokens@48 (k) | Q SD@48 | Credit@48 (%) |
|---|---|---|---|---|
| UCT (000) | ||||
| 001 | ||||
| 010 | ||||
| 011 | ||||
| 100 | ||||
| 101 |
| Evaluator | Mean utility | Tokens | Credit agreement (%) | |
|---|---|---|---|---|
| ASCT-Uniform | 0.5449 | 21,077.9 | 0.4571 | 76.7 |
| ASCT-UCT | 0.5887 | 20,788.5 | 0.5021 | 76.7 |
| ASCT-AgentUCT | 0.5960 | 20,531.3 | 0.5105 | 73.3 |