cs.LGSep 28, 2026

MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism

Authors: Tong Qiao, Ao Zhou, Yingjie Qi, Chunming Hu, Jianlei Yang

Organizations: School of Computer Science and Engineering, Beihang University, China · School of Software, Beihang University, China · State Key Laboratory of Complex and Critical Software Environment, Beihang University, China · Qingdao Research Institute, Beihang University, Qingdao, China

Abstract

Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated topology-aware bias matrix jointly incur significant per-layer memory overhead, and heavy graph embedding layers result in severe workload imbalances. These characteristics are unique to GT training and are not addressed by parallelism techniques designed for either conventional GNNs or Transformers, making a dedicated solution necessary. This paper introduces MegaGraph, the first automated hybrid parallelism framework designed for efficient GT training. MegaGraph designs three specialized strategies, namely graph-aware context parallelism, heterogeneous pipeline parallelism, and hybrid data parallelism, to support efficient training on large-scale graphs. However, coordinating these three parallelism strategies yields an exponentially large configuration space. To address this complexity, an automatic search engine leverages precise cost models via a Profile - Model - Search workflow to identify the optimal parallelism configuration. Evaluations demonstrate that MegaGraph enables training on large-scale graphs where state-of-the-art baselines fail due to out-of-memory (OOM) errors. The framework reduces per-device peak memory by up to 77.8% and achieves up to 4.51×\times training speedup while maintaining model accuracy.

Figures & tables

Explore similar work

CardsList
  1. Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs

    Apr 17, 2026Jun-Liang Lin, Kamesh Madduri, Mahmut Taylan KandemirMultiplex Graph TransformersCpu-Gpu Hybrid Designs

  2. On Efficient Scaling of GNNs via IO-Aware Layers Implementations

    May 29, 2026Daria Fomina, Daniil Krasylnikov, Alexey Boykov +3Graph Neural NetworksGraphics Processing Unit Kernels

  3. FAST: A Holistic Framework for Optimizing Memory-I/O, Computation, and Sampling in Temporal GNN Training

    Jul 6, 2026Yushu Cai, Qingrui Zhu, Lei Liu +3Temporal Graph Neural NetworksCpu-Gpu Hybrid Designs