cs.CRSep 29, 2026

Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks

Authors: Ying Song, Xiaowei Jia, Balaji Palanisamy

Organizations: University of Pittsburgh · Rutgers University

Abstract

As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model's black-box API, an adversary can construct a functionally equivalent surrogate model, compromising proprietary intellectual property and downstream security. Existing GNN stealing attacks, however, rely on overly permissive assumptions, such as soft-label outputs, large query budgets, full-graph query access, and prior knowledge of victim backbones that rarely hold in real-world deployments. In this work, we formalize a strictly constrained black-box, hard-label and backbone-agnostic threat model for GNN stealing attacks under a tight query budget. Given these realistic restrictions, we identify four fundamental challenges: sparse local structures and isolated nodes that degrade victim label quality, insufficient supervision signals, systematic imbalance with incomplete class coverage, and backbone mismatch. To address these interlocking barriers, we propose Dagger, a novel two-phase decoupling-based attack framework. Specifically, in Phase 1, Dagger pre-trains a surrogate using decoupled information propagation to preserve structural context over sparse local subgraphs while handling isolated nodes, combined with manifold-level node mixup to synthesize continuous supervision signals and smooth decision boundaries. In Phase 2, Dagger freezes the encoder and fine-tunes the classifier head via class-balanced sampling paired with logit adjustment to rectify severe query imbalance without requiring extra victim queries. Extensive experiments across four benchmark graphs and four GNN backbones demonstrate that Dagger consistently outperforms state-of-the-art GNN stealing attacks, achieving up to 18.16% higher fidelity while only utilizing 12.23×\times fewer queries than the strongest baseline.

Figures & tables

Explore similar work

CardsList
  1. GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It?

    May 12, 2026Kaixiang Zhao, Bolin Shen, Yuyang Dai +2Graph Neural NetworksWatermarking

  2. Blackknife: Hard-Label Query-Limited Black-Box Attacks on Heterogeneous Graph Neural Networks

    Jun 28, 2026Honglin Gao, Junhao Ren, Lan Zhao +3Graph Neural NetworksWhite-Box Spectral-Subspace-Guided Attack

  3. Defending against Model Extraction for GNNs with Model Reprogramming

    Aug 11, 2026Yan Wen, Zhenyi Wang, Heng HuangGraph Neural NetworksSemantic Firewall