cs.DBOct 8, 2026

Cost-Aware Mixture-of-Experts Coordination for Model Markets

Authors: Yizhou Ma, Wenbo Wu, Xikun Jiang, Zhuoqin Yang, Luis-Daniel Ibáñez

Organizations: University of Southampton Southampton, UK · Aalborg University Copenhagen, Denmark · University of Nottingham Ningbo, China

Abstract

Existing model marketplaces typically trade and select individual models as indivisible units, limiting their ability to exploit complementarities among heterogeneous experts. This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism. In this framework, brokers use gating networks to coordinate multiple heterogeneous experts and deliver a composite model service. We formalize the market participants, service workflow, expert cost structure, and a welfare objective that combines predictive utility with heterogeneous execution costs. We then derive a cost-aware gating mechanism and market-aware training objective, and introduce a cost-adjusted revenue allocation rule that distributes residual revenue according to realized expert participation and execution cost. We also establish basic theoretical properties of the allocation rule, including budget balance, participation monotonicity, and cost sensitivity. Experiments over five random seeds on fifteen tabular and image benchmarks use independently trained and frozen neural and tree-based experts together with latency-derived execution costs. MoE Market achieves the highest mean welfare on all fifteen datasets and a lower mean expected cost than Standard MoE in every case, while maintaining competitive predictive performance. The allocation experiments further demonstrate systematic sensitivity to expert participation and cost, together with substantially lower computational overhead than exact Shapley allocation. These results suggest that MoE can serve as a market-level coordination principle for collaborative, cost-aware, and economically grounded model marketplaces.

Figures & tables

Explore similar work

CardsList
  1. Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization

    Sep 30, 2026Zheng Lin, Shaoke Fang, Yuxin Zhang +6Submodular OptimizationMixture-of-Experts Pruning

  2. Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

    Aug 11, 2026Gongli Zhang, Zhulin Liu, C. L. Philip ChenMixture of ExpertsMixture-of-Experts Inference

  3. Beyond Task-Agnostic: Task-Aware Grouping for Communication-Efficient Multi-Task MoE Inference

    May 31, 2026Zhiyao Xu, Aoxue Liu, Zhanjie Ding +3Mixture-of-Experts InferenceMixture-of-Experts Models