Authors: Qiang Yang, Zhiqiang Kou, Xueyi Zhang, Dong-Dong Wu, Hanlin Gu, Jing Guo, Yang Liu, Di Jiang, +1 more
Organizations: The Hong Kong Polytechnic University, Hong Kong SAR, China · National University of Singapore, Singapore · The University of Tokyo, Japan · Hong Kong University of Science and Technology, Hong Kong SAR, China · Tengen AI, Hong Kong SAR, China
Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbf{Federated Agent Optimization (FAO)}, which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.
Figures & tables
Figure 1: One round of FAO. Clients upload typed updates Δk , the coordinator aggregates them into views Gk , and each client adapts its view. Utility is measured at Ak+ , leakage on Δk , and communication cost on every message that crosses the line.
Category
Shared update
Main leakage channel
Agent-level methods
Precursors
Policy ( Δπ )
adapter weights or gradients; soft or textual prompts
inversion of gradients; private examples kept in prompts
[ 186 , 187 , 185 , 240 ] , [ 35 , 29 ] ∗
[ 229 , 216 , 170 , 30 , 51 ]
Memory ( ΔM )
insights, reasoning abstractions, workflows
private content quoted in text
[ 212 , 109 ] ∗
[ 183 ] , [ 2 , 121 ] ∗
Tool use ( ΔT )
tool knowledge as text or typed fields; usage statistics
schemas of internal systems; arguments of calls
[ 200 , 25 ] ∗
—
Reward ( ΔR )
parameters of evaluators or preference models; scores
labels and outcomes of individual cases
—
[ 199 , 168 ] , [ 90 ] ∗
Structured knowledge ( ΔO )
embeddings of entities and relations; rules; skill patches
entity sets and triples; business processes
[ 75 , 209 ] ∗
[ 31 , 138 , 230 ]
Table 1: Taxonomy of FAO. Agent-level methods federate a component of a language agent, precursors the same kind of component outside an agent. ∗ Preprint.
Type
Shared artifact z
Example in identity verification
Federated methods
Experience
pattern distilled from successes and failures
normalize the transliteration of names before the registry lookup
[ 212 , 109 ]
Skill
parameterized procedure
recognize the license, query the registry, compare the legal representative, escalate on mismatch
[ 209 ]
Rule
effect of an action in the environment
three failed one-time codes lock the account for 30 minutes
[ 75 ]
Reward
evaluator or rubric
a dialogue is non-compliant if the agent reveals which field failed the check
[ 199 , 90 ] †
Ontology
alignment of concepts
beneficial owner ≡ ultimate controlling person
[ 31 , 138 ] †
Table 2: Five types of shared knowledge, which differ in function, with a running example. † Precursor that federates the component outside an agent (Table 1 ).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
K , k , X
number and index of clients; task space shared by all clients
Qk , uk(τ)
task distribution of client k ; utility of a trajectory
Dk , nk
private dataset of client k (tasks, logged trajectories, raw records); its number of tasks
Ak0 , Akt
initial agent of client k ; its agent after round t
parameters of the policy πk , continuous or discrete
Appendix
Table 3: Notation.
Method
Block
Shared update
Aggregation
Agent-level methods
FedMABench [ 186 ] , FedGUI [ 185 ]
Δπ
adapters of GUI agents based on vision-language models
standard federated algorithms
MobileA3gent [ 187 ]
Δπ
models trained on automatically annotated phone usage
averaging weighted by episode and step statistics
Fed-SE ∗ [ 35 ]
Δπ
adapters tuned on successful trajectories
unweighted averaging of adapters
FedAgent ∗ [ 29 ]
Δπ
parameters trained by reinforcement learning
federated averaging
FedWave [ 240 ]
Δπ
role-specific adapters of a multi-stage workflow
federated adapters fused by a shared mixture-of-experts router
Appendix
Table 4: Agent-level FAO methods and selected precursors, as instances of the typed update of Eq. ( 5 ). ∗ Preprint.
Shared type
Policy πk
Memory Mk
Tools Tk
Reward Rk
Structured knowledge Ok
Experience
demonstrations for prompting or fine-tuning
insights stored for retrieval
corrected tool descriptions, known failure cases
positive and negative examples for the evaluator
new rules derived from recurring patterns
Skill
procedures placed in the prompt
procedures retrieved on demand
compositions of tools
checkpoints that a correct trajectory must pass
extension of the skill library
Rule
preconditions of actions stated in the prompt
rules retrieved for similar states
expected effects and failure modes of tool calls
penalties for actions with known bad effects
extension of the rule base
Reward
training signal for policy optimization
filter for what is written to memory
reliability scores of tools
parameters or rubrics of the evaluator
validation of new rules and skills
Ontology
shared vocabulary in prompts
retrieval keys over aligned concepts
mapping of tool arguments to shared concepts
criteria stated over shared concepts
aligned concepts and relations
Appendix
Table 5: How knowledge of each type (rows) can improve each component of the receiving agent (columns). One shared artifact can improve several components.
College of Future Information Technology, Fudan University, Shanghai 200433, China · Division of Natural and Applied Sciences, Duke Kunshan University, Suzhou 215316, China · Department of Electrical and Computer Engineering, The University of British Columbia, BC V6T 1Z4, Canada +8