Agent Priors-guided Policy Learning
Organizations: National University of Singapore
Abstract
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.
Figures & tables
| Mean | |||||||||
| Method | IID | OOD | IID | OOD | IID | OOD | IID | OOD | OOD |
| Diffusion Policy (B0) | 73.33 | 28.96 | 82.50 | 37.29 | 97.50 | 38.54 | 100.00 | 46.67 | 37.86 |
| Relational prior (B1) | 85.83 | 37.92 | 93.33 | 42.71 | 100.00 | 52.29 | 100.00 | 56.25 | 47.29 |
| APPL, first proposal ( ) | 84.17 | 56.04 | 95.83 | 63.54 | 100.00 | 73.33 | 100.00 | 78.75 | 67.92 |
| APPL, best of three ( ) | 90.83 | 80.00 | 100.00 | 81.46 | 100.00 | 90.42 | 100.00 | 88.96 | 85.21 |
| APPL ( ) | 96.67 | 89.58 | 100.00 | 93.54 | 100.00 | 93.33 | 100.00 | 93.13 | 92.40 |
| Drawer | Buffer | Retrieve | Peg | Pour | Total (%) | ||||||||
| Method | M | T | M | T | M | T | M | T | M | T | M | T | Comp. |
| DP | 0/8 | – | 0/8 | – | 1/8 | – | 3/8 | – | 0/8 | – | 10.0 | – | – |
| SinglePrior | 1/8 | – | 0/8 | – | 2/8 | – | 1/8 | – | 0/8 | – | 10.0 | – | – |
| Agent+VLA | 1/8 | 8/8 | 1/8 | 4/8 | 5/8 | 8/8 | 5/8 | 4/8 | 1/8 | 4/8 | 32.5 | 70.0 | 6/16 |
| APPL w/o interface information | 0/8 | 4/8 | 4/8 | 5/8 | 2/8 | 6/8 | 0/8 | 7/8 | 2/8 | 4/8 | 20.0 | 65.0 | 2/16 |
| APPL w/o prior information | 4/8 | 7/8 | 0/8 | 4/8 | 4/8 | 6/8 | 3/8 | 6/8 | 7/8 | 6/8 | 45.0 | 72.5 | 3/16 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Task (native ID; ) | Factor | L | H | E-L | E-H |
|---|---|---|---|---|---|
| pick-place-wall | object | ||||
| (pick-place-wall-v3; 45) | goal | ||||
| assembly | nut | ||||
| (assembly-v3; 42) | peg goal | ||||
| drawer | cabinet | ||||
| (drawer-open-v3; 41) | cabinet yaw |
| Task | Success condition |
|---|---|
| pick-place-wall | Object-to-goal Euclidean distance . |
| assembly | RoundNut site’s horizontal target error and target height minus that site’s height . |
| drawer | for three consecutive control steps; joint validity requires throughout the episode. |
| door | Same progress, hold, and validity rule as drawer, with the adapted rotating-root geometry. |
| peg-insert-side | . |
| stick-push | Object2-to-target distance , and native grasp success: contact with the main object, positive aperture, and . |
| Setting | Value |
|---|---|
| History / horizon / executed actions | 2 observations / 16 / 4 |
| U-Net | Down dimensions ; step embedding 128; kernel 5; 8 GroupNorm groups |
| Optimizer | AdamW; learning rate (constant); ; weight decay |
| Updates / batch | 20,000 / 128; gradient-norm clipping 1.0; EMA decay .995 |
| Training diffusion | DDPM, 100 steps, epsilon prediction, squaredcos_cap_v2 schedule |
| Sampling | DDIM, 16 steps, |
| System | ||||
|---|---|---|---|---|
| B0 | 28.96 [26.88, 31.04] | 37.29 [35.00, 39.58] | 38.54 [36.25, 40.83] | 46.67 [44.17, 49.17] |
| B1 | 37.92 [35.83, 40.00] | 42.71 [41.04, 44.38] | 52.29 [50.42, 54.17] | 56.25 [55.42, 57.08] |
| 56.04 [53.96, 58.13] | 63.54 [61.66, 65.63] | 73.33 [71.46, 75.21] | 78.75 [77.29, 80.21] | |
| 80.00 [78.33, 81.67] | 81.46 [80.00, 82.92] | 90.42 [88.75, 91.88] | 88.96 [87.71, 90.21] | |
| 89.58 [87.92, 91.25] | 93.54 [91.88, 95.00] | 93.33 [92.08, 94.38] | 93.13 [91.46, 94.58] |
| Setting | Value |
|---|---|
| Data | 12 complete demonstrations per task; one normalizer per task |
| History / prediction / execution | 2 observations / 16 actions / 8 actions |
| Diffusion | 1D U-Net, 100 DDPM training and sampling steps |
| Optimization | AdamW; learning rate ; weight decay ; batch 128 |
| Schedule / gradient / EMA | Cosine, 500 warmup updates / norm limit 1 / decay 0.999; final EMA |
| Updates | 60,000 per DP and SinglePrior policy; 20,000 per APPL policy |
| Task | Skill (or full-task policy) | h01 | h02 | h03 | h04 (added) |
|---|---|---|---|---|---|
| Drawer exchange | open_drawer | 3/12 | 6/12 | 12/12 | 2/12 |
| red_transfer | 4/12 | 10/12 | 9/12 | 10/12 | |
| blue_insert | 10/12 | 12/12 | 10/12 | 11/12 | |
| full-task policies | DP 5/12, SinglePrior 5/12 | ||||
| Buffer exchange | buffer_red | 5/12 | 12/12 | 0/12 | 8/12 |
| place_blue | 7/12 | 1/12 | 12/12 | 4/12 | |
| Comparator | Suite | Wins | Losses | |
|---|---|---|---|---|
| DP | Motion OOD | 18 | 2 | |
| SinglePrior | Motion OOD | 17 | 1 | |
| w/o interface information | Motion OOD | 17 | 5 | 0.017 |
| w/o interface information | Task-level OOD | 12 | 1 | 0.003 |
| w/o interface information | Composition | 6 | 0 | 0.031 |
| w/o prior information | Motion OOD | 6 | 4 | 0.754 |
| Family | Shapes the policy | Informs the caller | Structure designed by |
|---|---|---|---|
| Structural priors | Yes | No; the prior remains in training | A human, per task family |
| Agents designing learning | Yes, one design per policy | No; the design serves training only | An agent |
| Generalist VLAs | Implicitly, through data | An instruction without stated scope | Data |
| Agents calling tools | No; tools are given | Names, feasibility estimates, or execution records | Designers or pretraining |
| Shared abstractions | Partly, through a fixed rule or recipe | Predicates, preconditions and effects, or skill specifications | A designer or a fixed rule |
| APPL | Yes, with alternative priors per skill | The same priors, with handoff overlap and training support | An agent, per skill |
| Method | Low-level policies | Structure designed by | Alternatives per skill | Reads | Decision maker |
|---|---|---|---|---|---|
| LGA ( Peng et al., 2024 ) | Imitation from few demonstrations | A language model, one state abstraction per task | No | Nothing; the abstraction is the policy input | None |
| KALM ( Fang et al., 2025 ) | Keypoint-conditioned imitation | A VLM proposes keypoints checked on demonstrations | No | Nothing | None |
| SymSkill ( Shao et al., 2025 ) | Dynamical-system skills in relative frames | Co-invented predicates; a VLM selects frames offline | No | Predicates and operators | Symbolic planner |
| Lorang et al. (2025) | Diffusion controllers | A fixed rule restricts inputs to operator-relevant objects | No | Learned operators | Symbolic planner |
| DR-LfD ( Chen et al., 2026 ) | Diffusion policies and equivariant primitives | A fixed rule based on contact complexity | No | Learned initiation and termination keyposes | Task and motion planner |
| MaestroMotif ( Klissarov et al., 2025 ) | Reinforcement-learned skills in NetHack | Human skill descriptions; LLM-derived rewards | No | Skill descriptions, initiation and termination code | LLM-written policy over skills |