cs.AIOct 4, 2026

Fusion is the New Mutation: Bandit-Guided Evolution on Workflow Graphs

Authors: Zhiwei Shang, Jiahang Sun, Mingrong Gong, Mingze Kong, Zikun Qu, Pingchen Lu, Junhao Dong, Zhipiao Liu, +4 more

Organizations: School of Data Science, The Chinese University of Hong Kong, Shenzhen, China · College of Computing and Data Science, Nanyang Technological University, Singapore · Huawei Technologies Co., Ltd., China · Artificial Intelligence Thrust, The Hong Kong University of Science and Technology (Guangzhou), China

Abstract

Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinations requires learning from limited fusion feedback. We introduce DAGO (Directed Acyclic Graph Optimization), a contextual-bandit-guided framework that learns which parent workflows to fuse under a limited evaluation budget. DAGO formulates each candidate parent combination as an arm, represented by pretrained embeddings of its constituent workflows' code and prompts. A diagonal LinUCB policy learns a shared reward model across arms and balances exploitation of arms with high predicted offspring quality against uncertainty-driven exploration. After an arm is selected, an LLM generates a child workflow through summary-guided fusion, and the child's validation score serves as the reward for updating the bandit. A shared directed acyclic graph maintains discovered workflows and their multi-parent lineage, providing an expanding pool of parents for subsequent arm proposals. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO achieves the highest macro-average score among the evaluated baselines. Under matched validation-evaluation budgets, it improves over AFlow from 80.3 to 81.7 while reducing aggregate search expenditure by 11.2%. Ablation studies show that LinUCB-guided arm selection outperforms both random selection and its exploration-free variant, supporting the value of feedback-driven selection and exploration-exploitation balance.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MoFlow: Multi-Objective Agentic Workflow Generation

    Sep 29, 2026Yining Lu, Aurelie Lozano, Xi Yang +3Agentic WorkflowsAgentic Workflow Design

  2. FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse

    Jun 9, 2026Lingzhi Yuan, Chenghao Deng, Fangxu Yu +3Agentic WorkflowsAgentic Workflow Design

  3. Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness

    Jul 27, 2026Yang Li, Hai Liu, Dian Shao +8Agentic WorkflowsAgentic Optimization