Federated Agent Optimization
Organizations: The Hong Kong Polytechnic University, Hong Kong SAR, China · National University of Singapore, Singapore · The University of Tokyo, Japan · Hong Kong University of Science and Technology, Hong Kong SAR, China · Tengen AI, Hong Kong SAR, China
Abstract
Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbf{Federated Agent Optimization (FAO)}, which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.
Figures & tables
| Category | Shared update | Main leakage channel | Agent-level methods | Precursors |
| Policy ( ) | adapter weights or gradients; soft or textual prompts | inversion of gradients; private examples kept in prompts | [ 186 , 187 , 185 , 240 ] , [ 35 , 29 ] ∗ | [ 229 , 216 , 170 , 30 , 51 ] |
| Memory ( ) | insights, reasoning abstractions, workflows | private content quoted in text | [ 212 , 109 ] ∗ | [ 183 ] , [ 2 , 121 ] ∗ |
| Tool use ( ) | tool knowledge as text or typed fields; usage statistics | schemas of internal systems; arguments of calls | [ 200 , 25 ] ∗ | — |
| Reward ( ) | parameters of evaluators or preference models; scores | labels and outcomes of individual cases | — | [ 199 , 168 ] , [ 90 ] ∗ |
| Structured knowledge ( ) | embeddings of entities and relations; rules; skill patches | entity sets and triples; business processes | [ 75 , 209 ] ∗ | [ 31 , 138 , 230 ] |
| Type | Shared artifact | Example in identity verification | Federated methods |
| Experience | pattern distilled from successes and failures | normalize the transliteration of names before the registry lookup | [ 212 , 109 ] |
| Skill | parameterized procedure | recognize the license, query the registry, compare the legal representative, escalate on mismatch | [ 209 ] |
| Rule | effect of an action in the environment | three failed one-time codes lock the account for 30 minutes | [ 75 ] |
| Reward | evaluator or rubric | a dialogue is non-compliant if the agent reveals which field failed the check | [ 199 , 90 ] † |
| Ontology | alignment of concepts | beneficial owner ultimate controlling person | [ 31 , 138 ] † |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
| , , | number and index of clients; task space shared by all clients |
| , | task distribution of client ; utility of a trajectory |
| , | private dataset of client (tasks, logged trajectories, raw records); its number of tasks |
| , | initial agent of client ; its agent after round |
| policy, memory, tools, reward, structured knowledge (ontologies, rules, skills) | |
| parameters of the policy , continuous or discrete |
| Method | Block | Shared update | Aggregation |
| Agent-level methods | |||
| FedMABench [ 186 ] , FedGUI [ 185 ] | adapters of GUI agents based on vision-language models | standard federated algorithms | |
| MobileA3gent [ 187 ] | models trained on automatically annotated phone usage | averaging weighted by episode and step statistics | |
| Fed-SE ∗ [ 35 ] | adapters tuned on successful trajectories | unweighted averaging of adapters | |
| FedAgent ∗ [ 29 ] | parameters trained by reinforcement learning | federated averaging | |
| FedWave [ 240 ] | role-specific adapters of a multi-stage workflow | federated adapters fused by a shared mixture-of-experts router | |
| Shared type | Policy | Memory | Tools | Reward | Structured knowledge |
| Experience | demonstrations for prompting or fine-tuning | insights stored for retrieval | corrected tool descriptions, known failure cases | positive and negative examples for the evaluator | new rules derived from recurring patterns |
| Skill | procedures placed in the prompt | procedures retrieved on demand | compositions of tools | checkpoints that a correct trajectory must pass | extension of the skill library |
| Rule | preconditions of actions stated in the prompt | rules retrieved for similar states | expected effects and failure modes of tool calls | penalties for actions with known bad effects | extension of the rule base |
| Reward | training signal for policy optimization | filter for what is written to memory | reliability scores of tools | parameters or rubrics of the evaluator | validation of new rules and skills |
| Ontology | shared vocabulary in prompts | retrieval keys over aligned concepts | mapping of tool arguments to shared concepts | criteria stated over shared concepts | aligned concepts and relations |