Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making
Organizations: Department of Electrical and Computer Engineering, Sungkyunkwan University, Republic of Korea. · Soongsil University, Republic of Korea.
Abstract
Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals--global entropy and local top-2 margin--computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the performance-cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.4% while maintaining a token budget comparable to ReAct, yielding a improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.
Figures & tables
| Model | Type | Method | ALFWorld | BabyAI | WebShop | |||||||||
| Succ | #Horizon | #Token | Eff | Succ | #Horizon | #Token | Eff | Succ | #Horizon | #Token | Eff | |||
| GPT-4.1 | Basic | ReAct | ||||||||||||
| Qwen2.5- 7B-Instruct | Basic | ReAct | ||||||||||||
| Deliberation | Reflexion | |||||||||||||
| Deliberation | Debate | |||||||||||||
| Deliberation | CER | |||||||||||||
| Global | Local | Variant | Succ | #Horizon | #Token |
| ✓ | ✓ | Full (SAG) | 78.3 | 11.4 | 56 |
| ✗ | ✓ | w/o | 77.6 | 12.1 | 51 |
| ✓ | ✗ | w/o | 74.6 | 11.4 | 52 |
| ✗ | ✗ | Always (no gating) | 62.0 | 17.4 | 282 |
| ✗ | ✗ | w/o critic | 33.0 | 19.3 | 50 |
| Method | Critique | Succ. | #Token |
| ReAct + online SFT | None | 30.2 | 57 |
| LAC + online SFT | Always-on | 74.2 | 341 |
| SAG | Gated | 78.3 | 56 |
| Appendix: Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Description |
| Time step in an interaction episode | |
| Environment observation at step | |
| Actor context at step , including task instruction, observation, and truncated history | |
| Set of admissible actions at step | |
| Action executed at step | |
| Actor reasoning state generated from at step |
| [Actor Prompt] You are an agent for household tasks. Your goal is to complete the task described in natural language by reasoning step-by-step and producing an action. Always output in the format: "think: reasoning here" "action: [chosen action]" Actions must come from the admissible set provided. Do not invent new actions. [Critic Prompt] You are a critic for an LLM-based agent. Your GOAL: detect and correct hallucinations, loops, or inconsistent actions. INPUTS: - task: the goal - history: what happened so far - actions: valid moves. INSTRUCTIONS: 1. verify objects: does the object the agent is targeting MATCH the task description? 2. verify history: did the agent already check this empty location or repeat failures? 3. decision: - If the agent is safe and correct, output "next_action": "continue". - If the agent is wrong, provide a specific "next_action" to fix it. 4. reasoning: - Also provide a short explanation in "reasoning". - The reasoning should explain why the proposed action is appropriate. - Keep it concise, one or two sentences. |
| [Actor Prompt] You are an agent in the Grid world. Your goal is to follow the natural language instruction step-by-step. At each step, reason explicitly before acting. Use the format: "think: reasoning here" "action: [chosen action]" Primitive actions include: go left, go right, go forward, pick up object, drop object, etc. [Critic Prompt] You are a critic helping an actor agent navigate a small grid world. Your GOAL: - Detect when the actor is stuck, uncertain, or moving inefficiently. - Propose a corrective plan that helps the actor make progress toward the goal. INPUTS: - goal: The current task the actor must complete. - observation: The actor’s current state in the grid world. - valid actions: The list of actions the actor can execute now. INSTRUCTIONS: 1. check progress: - Determine whether the actor is making progress toward the goal. - If the actor seems stuck in a loop, suggest actions that break the loop. 2. check validity: - Only use actions from the valid actions list. - Do not invent new actions. 3. proposed plan: - Propose a SHORT sequence of 1 to 3 actions. - The plan should be directly executable in the current state. 4. reasoning: - Provide a brief reason explaining why the plan is helpful. - Keep the reasoning concise, at most one sentence. |
| [Actor Prompt] You are a shopping assistant. Your task is to satisfy the given shopping goal. At each step, reason aloud and then select one action. Use search[query] for the initial query; after search results are shown, choose from the web actions available in the current observation, such as click[item], click[option], navigation, or click[Buy Now]. Use format: "think: reasoning here" "action: [web action]" [Critic Prompt] You are a critic for a WebShop shopping agent. GOAL: - Help the actor choose the best next action for completing the shopping task. - Detect inefficient actions, loops, or choices that do not satisfy the user instruction. INPUTS: - user instruction: the shopping request, including constraints such as product type, size, flavor, count, price, rating, or other preferences. - recent history: The recent actions and observations from the shopping trajectory. - candidate actions: The list of actions currently available to the actor. INSTRUCTIONS: 1. check constraints: - Identify the important constraints in the user instruction. - Prefer actions that move toward satisfying all constraints. 2. check progress: - Avoid repeating unhelpful actions. - Avoid going back unless it clearly helps recover from a mistake. - Prefer clicking the most relevant product, option, or navigation action. 3. choose action: - Choose exactly ONE best next action. - The action must be one of the candidate actions. - If no candidate action is clearly useful, choose the safest progress-making action. - Use "none" only when no valid progress-making action exists. 4. rationale: - Provide a very short reason explaining why the chosen action is appropriate. - Keep the rationale concise, one sentence at most. |
| Component | Value | Component | Value |
| Uncertainty-based critique gate | |||
| Gate threshold | 0.4 | Margin threshold | 0.2 |
| Online SFT (self-improvement) | |||
| History window | 8 steps | Max obs. length | 800 chars |
| Buffer capacity | 50,000 | Successes per update | 10 |
| Train batch size | 2 | Grad. accumulation | 8 |
| Dimension | RL-based agent post-training | SAG |
| Primary goal | Improve the underlying policy before deployment | Decide when to use critique during deployment |
| Typical methods | PPO, GRPO, GiGPO | Gated critique + online SFT |
| Training signal | Rewards, preferences, or advantage estimates | Successful trajectories from online interaction |
| Rollout structure | Multiple sampled trajectories/completions per task | One interaction stream with selective critique |
| Models during training | Actor, often reference policy, optimizer states | Actor update; critic fixed and queried selectively |
| Cost focus | Training-time rollout and optimization budget | Inference-time token, critic-call, and latency cost |
| Benchmark | Backbone | w/o critic | always-on critic | SAG | ||||||
| Succ. | Token | Call | Succ. | Token | Call | Succ. | Token | Call | ||
| ALFWorld | 3B | 21.6 | 57 | – | 59.0 | 271 | 15.1 | 69.4 | 58 | 3.2 |
| 7B | 25.4 | 55 | – | 74.8 | 273 | 14.7 | 78.3 | 56 | 2.6 | |
| 14B | 77.2 | 49 | – | 80.5 | 292 | 13.3 | 82.1 | 53 | 2.3 | |
| 32B | 79.1 | 48 | – | 86.2 | 308 | 13.2 | 87.2 | 54 | 2.1 | |
| BabyAI | 3B | 34.0 | 59 | – | 52.3 | 356 | 10.3 | 61.8 | 79 | 3.3 |
| Actor / Critic | Succ. |
| 7B / No critic | 25.4 |
| 7B / 3B critic | 76.9 |
| 7B / 7B critic | 78.3 |
| 7B / 14B critic | 82.1 |
| Method | ALFWorld | BabyAI-Text | WebShop | |||
| Succ. | #Token | Succ. | #Token | Succ. | #Token | |
| ReAct + online SFT | 30.2 | 57 | 40.8 | 56 | 15.3 | 12 |
| LAC + online SFT | 74.2 | 341 | 67.2 | 438 | 32.1 | 84 |
| SAG | 78.3 | 56 | 70.0 | 78 | 37.0 | 9 |
| Order Setting | Seed / Order | Success (%) | Token (k) |
| Default benchmark order | default | ||
| Shuffled order | 1 | ||
| Shuffled order | 2 | ||
| Mean std over shuffled orders | default/1/2 | 78.3 0.3 | 56.7 1.2 |
| Succ. | #Token | Succ. | #Token | ||
| 0.2 | 79.0 | 74 | 0.1 | 77.1 | 53 |
| 0.3 | 78.9 | 63 | 0.2 | 78.3 | 56 |
| 0.4 | 78.3 | 56 | 0.4 | 78.1 | 61 |
| 0.5 | 75.1 | 49 | 0.5 | 77.8 | 64 |
| 0.7 | 66.1 | 45 | 1.0 | 75.3 | 60 |
| Method | ALFWorld | BabyAI-Text | WebShop | |||
| Succ. | #Token | Succ. | #Token | Succ. | #Token | |
| One-shot SFT | 72.8 | 63 | 66.9 | 82 | 34.8 | 10 |
| Online SFT (SAG) | 78.3 | 56 | 70.0 | 78 | 37.0 | 9 |
| Successes / Update | Succ. | Critic Calls / Ep. | #Token |
| 5 | 76.9 | 2.52 | 58 |
| 10 | 78.3 | 2.10 | 56 |
| 20 | 78.1 | 2.28 | 57 |
| 30 | 77.7 | 2.42 | 58 |
| Prompt variant | Step | Rationale by critic |
| Full prompt | 5 | Search shelves for both bowl and desklamp |
| w/o expected JSON example | 5 | Search shelves as likely bowl location |
| w/o goal | 5 | Bowl may be on an unchecked shelf |
| Full prompt | 9 | Check whether bowl is near the found desklamp |
| w/o expected JSON example | 9 | Re-examine shelf for missed items |
| w/o goal | 9 | Bowl may be near the desklamp |
| ID | OOD | |||||
| SFT target | Actor SR | SAG SR | Calls | Actor SR | SAG SR | Calls |
| No SFT | 24.6 | 72.3 | 3.00 | 25.7 | 70.7 | 2.85 |
| Matched reasoning | 43.2 | 76.0 | 1.35 | 42.9 | 77.8 | 1.67 |
| Action-only | 45.7 | 80.3 | 1.18 | 43.4 | 80.9 | 1.41 |
| Score quartile | Q1 | Q2 | Q3 | Q4 |
| Number of states | 45 | 45 | 44 | 44 |
| Disagreements | 8 | 10 | 14 | 20 |
| Disagreement rate (%) | 17.8 | 22.2 | 31.8 | 45.5 |
| You are in the middle of a room. Looking quickly around you, you see an armchair 1, a cabinet 4, a cabinet 3, a cabinet 2, a cabinet 1, a drawer 5, a drawer 4, a drawer 3, a drawer 2, a drawer 1, a dresser 1, a garbagecan 1, a safe 1, a shelf 12, a shelf 11, a shelf 10, a shelf 9, a shelf 8, a shelf 7, a shelf 6, a shelf 5, a shelf 4, a shelf 3, a shelf 2, a shelf 1, a sidetable 1, and a sofa 1. Your task is to: put two pillow in sofa. > [Agent] Act 1: think: To solve the task, I need to find and take the first pillow, then put it in sofa, then find and take the second pillow, then put it in sofa. OK. > [Agent] Act 2: think: First I need to find the first pillow. A pillow is more likely to appear in armchair (1), cabinet (1-4), drawer (1-5), dresser (1), sidetable (1), sofa (1). I can check one by one, starting with armchair 1. OK. > [Agent] Act 3: go to armchair 1 You arrive at armchair 1. On the armchair 1, you see a cellphone 2, a pillow 2, and a pillow 1. > [Agent] Act 4: take pillow 1 from armchair 1 You pick up the pillow 1 from the armchair 1. > [Agent] Act 5: think: Now I find the first pillow (1). Next, I need to put it in sofa 1. OK. > [Agent] Act 6: go to sofa 1 You arrive at sofa 1. On the sofa 1, you see a newspaper 3. > [Gate] global ambiguity , local margin . > [Gate] score (VoI proxy cost) critic invoked. > [Critic] Suggested action: move pillow 1 to sofa 1. > [Critic] Rationale: You are close, but you forgot to put the first pillow in the sofa before searching for the second one. Putting the first pillow in the sofa now reduces the risk of forgetting it later. > [Agent] Act 7: move pillow 1 to sofa 1 On the sofa 1, you see a newspaper 3, and a pillow 1. > [Agent] Act 8: go to armchair 1 You arrive at armchair 1. On the armchair 1, you see a cellphone 2, and a pillow 2. > [Agent] Act 9: take pillow 2 from armchair 1 You pick up the pillow 2 from the armchair 1. > [Agent] Act 10: think: Now I find the second pillow (2). Next, I need to put it in sofa 1. OK. > [Agent] Act 11: go to sofa 1 You arrive at sofa 1. On the sofa 1, you see a newspaper 3, and a pillow 1. > [Agent] Act 12: move pillow 2 to sofa 1 You move the pillow 2 to the sofa 1. Outcome: success > [SAG] Successful episode. Trajectory is added to the self-improvement buffer. |
| You are in the middle of a room. Looking quickly around you, you see a cabinet 4, a cabinet 3, a cabinet 2, a cabinet 1, a countertop 1, a garbagecan 1, a handtowelholder 2, a handtowelholder 1, a sinkbasin 2, a sinkbasin 1, a toilet 1, a toiletpaperhanger 1, and a towelholder 1. Your task is to: put a clean cloth in countertop. > [Agent] Act 1: think: To solve the task, I need to find and take a cloth, clean it with a sinkbasin, then place it on the countertop. OK. > [Agent] Act 2: think: First I need to find a cloth. I will check the cabinets starting from cabinet 4. OK. > [Agent] Act 3: go to cabinet 4 You arrive at cabinet 4. The cabinet 4 is closed. > [Agent] Act 4: open cabinet 4 You open the cabinet 4. The cabinet 4 is open. In it, you see a soapbottle 2. > [Agent] Act 5: go to cabinet 3 You arrive at cabinet 3. The cabinet 3 is closed. > [Agent] Act 6: open cabinet 3 You open the cabinet 3. The cabinet 3 is open. In it, you see a cloth 2. > [Agent] Act 7: take cloth 2 from cabinet 3 You pick up the cloth 2 from the cabinet 3. > [Agent] Act 8: go to sinkbasin 1 You arrive at sinkbasin 1. On the sinkbasin 1, you see a soapbar 1. > [Agent] Act 9: clean cloth 2 with sinkbasin 1 You clean the cloth 2 using the sinkbasin 1. > [Agent] Act 10: go to countertop 1 You arrive at countertop 1. On the countertop 1, you see a candle 1, a cloth 3, and a spraybottle 1. > [Agent] Act 11: think: I might need to fetch soap again before placing the cloth. I should go back to cabinet 4. OK. > [Gate] global ambiguity , local margin . > [Gate] score (VoI proxy cost) critic invoked. > [Critic] Suggested action: move cloth 2 to countertop 1. > [Critic] Rationale: The cloth you are holding is already clean. The task only requires placing a clean cloth on the countertop. Fetching soap again is unnecessary and risks delaying completion. > [Agent] Act 12: move cloth 2 to countertop 1 On the countertop 1, you see a candle 1, a cloth 3, and a spraybottle 1. > [Agent] Act 13: move cloth 2 to countertop 1 You move the cloth 2 to the countertop 1. Outcome: success > [SAG] Successful episode. The trajectory is added to the self-improvement buffer . |
| You are in the middle of a room. Looking quickly around you, you see a cabinet 6, a cabinet 5, a cabinet 4, a cabinet 3, a cabinet 2, a cabinet 1, a coffeemachine 1, a countertop 3, a countertop 2, a countertop 1, a drawer 3, a drawer 2, a drawer 1, a fridge 1, a garbagecan 1, a microwave 1, a shelf 3, a shelf 2, a shelf 1, a sinkbasin 1, a stoveburner 4, a stoveburner 3, a stoveburner 2, a stoveburner 1, and a toaster 1. Your task is to: put a hot apple in fridge. > [Agent] Act 1: think: First I need to find an apple. I will start by checking the fridge. OK. > [Agent] Act 2: go to fridge 1 You arrive at fridge 1. The fridge 1 is closed. > [Agent] Act 3: open fridge 1 You open the fridge 1. In it, you see a bowl 1, a cup 2, a cup 1, a lettuce 1, a mug 2, a potato 1, and a tomato 1. > [Agent] Act 4: think: No apple here. I might need to search cabinets one by one. OK. > [Gate] global ambiguity , local margin . > [Gate] score (VoI proxy cost) critic invoked. > [Critic] Suggested action: go to countertop 1. > [Critic] Rationale: Apples are more likely to be placed on open surfaces such as countertops. Searching cabinets first is less efficient. > [Agent] Act 5: go to countertop 1 You arrive at countertop 1. On the countertop 1, you see an apple 1, an apple 2, and a mug 3. > [Agent] Act 6: take apple 1 from countertop 1 You pick up the apple 1 from the countertop 1. > [Agent] Act 7: go to microwave 1 You arrive at microwave 1. The microwave 1 is closed. > [Agent] Act 8: heat apple 1 with microwave 1 You heat the apple 1 using the microwave 1. > [Agent] Act 9: think: The apple is now hot. I should place it in the fridge. OK. > [Agent] Act 10: go to fridge 1 You arrive at fridge 1. The fridge 1 is open. > [Agent] Act 11: move apple 1 to fridge 1 You move the apple 1 to the fridge 1. Outcome: success > [SAG] Successful trajectory is added to the self-improvement buffer . |
| Your mission: pick up the green key. You start in a corridor with a blue key on the floor and a closed door in front. > [Agent] Act 1: think: The green key is not visible. I should move toward the door. OK. > [Agent] Act 2: go forward You are now standing directly in front of the door. > [Gate] global ambiguity , local margin . > [Gate] score critic invoked. > [Critic] Suggested action: open the door. > [Critic] Rationale: Progress requires entering the next room; waiting in front of the door adds no information. > [Agent] Act 3: open door The door opens. Inside you see a red ball in front, a green key on the right, and a yellow box. > [Agent] Act 4: go forward > [Agent] Act 5: turn left > [Agent] Act 6: turn right > [Gate] global ambiguity , local margin . > [Gate] score critic invoked. > [Critic] Suggested action: turn right toward the green key. > [Critic] Rationale: The goal object is already identified; reorientation loops are unnecessary. > [Agent] Act 7: turn right > [Agent] Act 8: go forward > [Agent] Act 9: pick up green key > [Agent] Act 10: stop Outcome: success > [SAG] Successful trajectory with two gated interventions is added to the self-improvement buffer . |
| Your task: purchase a red T-shirt under 15), red T-shirt (18), green hoodie (10). > [Agent] Act 2: think: There are two red T-shirts. I might choose the \bar{H}{t}=0.33m{t}=0.04<\deltaS_{t}=1.33>\tau\Rightarrow18 red T-shirt to the cart. > [Critic] Rationale: The 18) to cart > [Agent] Act 4: buy Outcome: Reward 1.0, success > [SAG] Successful trajectory is added to self-improvement buffer. |
| Your task: purchase a cheap black backpack. > [Agent] Act 1: click "bags" The bags page shows: black backpack (25), red backpack (55 one might be better quality. > [Gate] global ambiguity , local margin . > [Gate] score critic invoked. > [Critic] Suggested action: add the 55 violates this constraint. > [Agent] Act 3: add black backpack (\bar{H}{t}=0.31m{t}=0.06<\deltaS_{t}=1.31>\tau\Rightarrow$ critic invoked. > [Critic] Suggested action: place the order. > [Critic] Rationale: No further browsing is required once all constraints are satisfied. > [Agent] Act 7: buy Outcome: Reward 1.0, success |