We propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a form of self-prompting, GCR enables the model to refine its internal structure using its own outputs. While deep networks learn rich representations, these often capture noisy inter-class similarities that contradict the model's predicted semantics. GCR addresses this issue by introducing parameter-free Graph Consistency Layers (GCLs) at arbitrary depths. Each GCL builds a batch-level feature similarity graph and aligns it with a global, class-aware masked prediction graph, derived by modulating softmax prediction similarities with intra-class indicators. This alignment enforces that feature-level relationships reflect class-consistent prediction behavior, acting as a semantic regularizer throughout the network. Unlike prior work, GCR introduces a multi-layer, cross-space graph alignment mechanism with adaptive weighting, where layer importance is learned from graph discrepancy magnitudes. This allows the model to prioritize semantically reliable layers and suppress noisy ones, enhancing feature quality without modifying the architecture or training procedure. GCR is model-agnostic, lightweight, and improves semantic structure across various networks and datasets. Experiments show that GCR promotes cleaner feature structure, stronger intra-class cohesion, and improved generalization, offering a new perspective on learning from prediction structure. Project websiteCode
Figures & tables
Figure 1 : Relational graph visualization using a batch of 64 samples from CIFAR-10 on (left) DenseNet-121 and (right) MobileNet. We compare the baselines with their counterparts augmented by our GCLs. Our method promotes richer, class-aware semantic representations by acting as a form of self-prompting . For DenseNet-121, the baseline feature relational graph tends to connect samples based on superficial visual similarity ( e.g . , deer, horse, and automobile), often ignoring semantic boundaries. In contrast, our GCL-enhanced model produces more semantically coherent groupings , clearly separating animals from vehicles. On MobileNet, the prediction relational graph further highlights the strength of our method, demonstrating cleaner, more distinct class relationships compared to the baseline. These improvements reflect the effectiveness of our model in aligning feature and prediction spaces with semantic structure, despite being lightweight and parameter-free .
Figure 2 : Our parameter-free Graph Consistency Layer (GCL), highlighted in red, can be inserted after any micro-network block ( e.g . , Inception) or specific layer ( e.g . , fully connected). Each GCL constructs a relational graph from batch-level features using a similarity metric ( e.g . , cosine). A reference graph is generated from softmax predictions and masked by intra-class indicators: binary masks identifying semantically consistent pairs. Each GCL enforces alignment between masked prediction graph and the feature-level graphs. The resulting consistency signals are adaptively weighted, forming the Graph Consistency Regularization (GCR) framework, which integrates with the primary loss ( e.g . , cross-entropy), acting as a semantic regularizer to guide learning.
Figure 3 : Feature map visualizations from models trained on identical data batches : (top) baseline and (bottom) our GCL-augmented model. Brighter red regions indicate stronger feature activations. Compared to the baseline, GCL-enhanced maps more clearly emphasize class-discriminative cues, e.g . , cat faces, ears, and eyes, and for dogs, tongues, noses, and facial contours, reflecting improved focus and interpretability. GCL also yields higher classification accuracy (98.1% → 99.8%).
Figure 4 : Relational graph visualization on Kaggle cats vs . dogs. We compare the best baseline model and our GCL-augmented model using the same batch of 32 samples (red = cat, blue = dog). The baseline consists of four convolutional blocks and two fully connected layers; our method inserts a Graph Consistency Layer (GCL) after each, totaling six GCLs. The top row shows the baseline (without GCLs); the bottom row shows our GCL-enhanced model. Each column visualizes the relational graph at a specific layer, from early features (left) to final predictions (right). Early layers exhibit weak connectivity, as low-level features poorly capture class semantics. As depth increases, both models shift toward more structured, class-separable relationships. GCLs amplify this effect by attenuating low-similarity inter-class edges and reinforcing intra-class coherence, leading to improved accuracy (98.1% vs . 99.8%). For clarity, edges with similarity < 0.4 are omitted.
MAE
MNet
SN
SQNet
GLNet
Rx-50
Rx-101
R34
R50
R101
D121
Mean
Baseline
88.95 ±0.33
90.23 ±0.25
91.21 ±0.28
92.30 ±0.25
94.10 ±0.26
94.57 ±0.29
95.12 ±0.30
94.83 ±0.25
95.03 ±0.28
95.22 ±0.31
95.01 ±0.27
93.32 ±2.26
Early GCL
89.42 ±0.25
91.17 ±0.22
92.33 ±0.33
92.59 ±0.21
94.89 ±0.23
95.48 ±0.22
95.63 ±0.25
95.55 ±0.18
95.57 ±0.23
95.39 ±0.26
95.81 ±0.17
93.98 ±2.22
Mid GCL
89.77 ±0.22
91.15 ±0.18
92.58 ±0.19
92.40 ±0.20
94.82 ±0.21
95.47 ±0.19
95.39 ±0.24
95.69 ±0.23
95.61 ±0.20
95.75 ±0.17
95.51 ±0.22
94.01 ±2.15
Late GCL
89.70 ±0.29
91.40 ±0.19
92.36 ±0.21
92.80 ±0.19
94.88 ±0.19
95.35 ±0.28
95.71 ±0.26
95.69 ±0.19
95.66 ±0.17
95.51 ±0.24
95.72 ±0.22
94.07 ±2.14
Early+Mid
89.52 ±0.19
90.77 ±0.26
92.56 ±0.21
92.27 ±0.25
94.79 ±0.18
95.33 ±0.27
95.55 ±0.23
95.46 ±0.20
95.51 ±0.21
95.37 ±0.19
95.64 ±0.20
93.89 ±2.22
Mid+Late
89.59 ±0.28
91.23 ±0.20
92.79 ±0.20
92.86 ±0.23
94.61 ±0.22
95.51 ±0.19
95.38 ±0.27
95.45 ±0.18
95.33 ±0.26
95.52 ±0.14
95.70 ±0.19
94.00 ±2.09
Table 1 : Accuracy (%) on CIFAR-10 across models. Results are shown for MobileNet (MNet), ShuffleNet (SN), SqueezeNet (SQNet), GoogLeNet (GLNet), ResNeXt-50/101 (Rx-50/101), ResNet-34/50/101 (R34/R50/R101), DenseNet-121 (D121), and MAE under various GCL configurations (Early, Mid, Late, combinations, Full). Bold indicates the best improvements over baselines; underlines mark the second-best. The final column shows the average accuracy for each configuration.
MAE
MNet
SN
SQNet
Rx-50
Rx-101
R34
R50
D121
Mean
Baseline
64.29 ±0.34
65.95 ±0.25
70.11 ±0.30
69.43 ±0.27
77.75 ±0.29
77.83 ±0.30
76.82 ±0.28
77.31 ±0.29
77.09 ±0.27
72.95 ±5.50
Early GCL
65.05 ±0.29
67.45 ±0.21
71.96 ±0.27
70.90 ±0.20
79.18 ±0.22
79.69 ±0.27
77.90 ±0.22
79.37 ±0.25
79.41 ±0.22
74.55 ±5.78
Mid GCL
64.99 ±0.30
67.88 ±0.21
71.89 ±0.24
70.21 ±0.25
79.07 ±0.19
79.28 ±0.26
77.83 ±0.20
78.90 ±0.24
79.26 ±0.21
74.37 ±5.66
Late GCL
65.54 ±0.27
68.32 ±0.20
71.42 ±0.24
70.55 ±0.22
79.54 ±0.20
79.83 ±0.21
78.31 ±0.20
79.42 ±0.21
79.69 ±0.23
74.74 ±5.73
Early+Mid
65.23 ±0.31
67.62 ±0.24
71.50 ±0.28
70.47 ±0.19
78.90 ±0.18
79.25 ±0.20
77.41 ±0.19
78.58 ±0.24
79.22 ±0.20
74.28 ±5.56
Mid+Late
65.27 ±0.28
68.33 ±0.19
71.63 ±0.28
70.30 ±0.22
78.91 ±0.17
79.57 ±0.21
77.30 ±0.20
78.85 ±0.22
79.54 ±0.24
74.41 ±5.55
Table 2 : Accuracy (%) on CIFAR-100 across models.
Figure 5 : Relational graph comparison across five models on the same batch. Top row: baselines; bottom row: GCL-augmented versions, showing sparser inter-class connections and stronger class-aware structure , highlighting GCL’s effectiveness in enhancing relational representations.
ViT /32
ViT /16
CeiT
MViT XXS
MViT XS
MViT
Swin
MNet
R18SD
SER18
R34
Mean
Baseline
37.79 ±0.35
40.05 ±0.33
49.95 ±0.29
49.28 ±0.29
51.58 ±0.27
52.68 ±0.27
54.27 ±0.25
57.81 ±0.25
63.49 ±0.26
65.65 ±0.24
67.51 ±0.25
53.64 ±9.62
Early GCL
39.02 ±0.29
40.98 ±0.19
51.22 ±0.20
50.11 ±0.28
51.33 ±0.26
53.91 ±0.22
54.88 ±0.25
57.93 ±0.21
63.81 ±0.19
66.52 ±0.22
67.79 ±0.19
54.32 ±9.39
Mid GCL
38.61 ±0.23
40.95 ±0.19
50.30 ±0.19
49.92 ±0.26
51.43 ±0.22
53.88 ±0.20
55.23 ±0.24
57.63 ±0.20
64.03 ±0.22
65.66 ±0.23
67.62 ±0.20
54.11 ±9.38
Late GCL
37.98 ±0.28
40.35 ±0.25
50.82 ±0.20
49.77 ±0.21
51.99 ±0.23
54.10 ±0.19
55.47 ±0.21
57.87 ±0.23
63.79 ±0.19
65.85 ±0.25
67.61 ±0.19
54.18 ±9.56
Early+Mid
39.08 ±0.25
41.26 ±0.18
50.25 ±0.25
49.73 ±0.22
51.57 ±0.19
53.91 ±0.23
54.95 ±0.19
57.49 ±0.19
64.18 ±0.20
65.86 ±0.24
67.74 ±0.23
54.18 ±9.32
Mid+Late
38.44 ±0.18
40.52 ±0.28
50.09 ±0.25
50.55 ±0.18
51.48 ±0.21
53.90 ±0.20
55.62 ±0.23
57.65 ±0.21
64.29 ±0.17
65.95 ±0.19
67.58 ±0.21
54.19 ±9.52
Table 3 : Accuracy (%) on Tiny ImageNet across models. All results are obtained by training models from scratch. We also evaluate Stochastic ResNet-18 (R18SD) and SE-ResNet-18 (SER18).
Figure 6 : Performance gains ( Δ ,%) from adding GCLs with different configurations across various weighting schemes. Darker red indicates larger gains. Adaptive weighting achieves the highest improvement, showing the value of graph misalignment as a guidance signal. (c) and (d) show that GCL-augmented ShuffleNet yields more compact intra-class clusters and better inter-class separation.
Method
iFormer-S
iFormer-B
ViT-B/16
ViG-B
Baseline
83.4 ±0.40
84.6 ±0.45
74.3 ±0.51
82.3 ±0.42
Early GCL
83.8 ±0.31
85.0 ±0.40
74.7 ±0.44
82.8 ±0.35
Mid GCL
83.8 ±0.39
85.5 ±0.33
75.2 ±0.36
83.0 ±0.34
Late GCL
84.5 ±0.29
86.1 ±0.30
75.8 ±0.33
84.0 ±0.30
Early + Mid
84.3 ±0.33
85.9 ±0.38
75.6 ±0.41
83.7 ±0.33
Mid + Late
84.8 ±0.28
85.9 ±0.37
75.6 ±0.34
83.9 ±0.30
Table 4 : Comparison of iFormer, ViT, and ViG with different GCL integration strategies on ImageNet-1K. Results are averaged over three runs.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : t-SNE visualizations of feature representations on CIFAR-10. (Left) Original model architectures; (Right) corresponding GCL-augmented models. Our method consistently enhances feature structure across diverse architectures, including MobileNet, ResNet-34, and MAE (with ViT-Tiny encoder), by yielding tighter intra-class clusters and improved inter-class separation. Notably, GCR distinctly separates semantic groups such as animals and vehicles ( e.g . , (e) vs . (f)). Importantly, GCL is lightweight and introduces no additional parameters.
Figure 8 : Effect of batch size ( n ) in the GCR framework. We evaluate using the Masked Autoencoder model. The top row shows the baseline; the bottom row shows the GCL-augmented counterpart. From left to right, the relational graphs are constructed on softmax predictions as n increases from 16 to 128. As batch size grows, GCL-augmented models consistently exhibit tighter intra-class clusters and clearer inter-class separation.
ShuffleNet on CIFAR-10
CeiT on Tiny ImageNet
Batch Size
+ GCR
Baseline
Batch Size
+ GCR
Baseline
16
79.90 ± 0.38
78.88 ± 0.41
16
44.84 ± 0.31
43.78 ± 0.35
32
87.91 ± 0.36
86.91 ± 0.37
32
47.55 ± 0.29
46.89 ± 0.31
64
91.26 ± 0.25
90.64 ± 0.35
64
49.19 ± 0.25
48.09 ± 0.30
128
92.79 ± 0.20
91.21 ± 0.28
128
51.22 ± 0.20
49.95 ± 0.29
256
92.89 ± 0.25
92.07 ± 0.27
256
50.77 ± 0.19
49.62 ± 0.24
Appendix
Table 5 : Effect of batch size on GCR performance for ShuffleNet (CIFAR-10) and CeiT (Tiny ImageNet).
Method
Cosine
RBF
Polynomial
Sigmoid
Laplacian
Baseline
65.95 ± 0.25
65.95 ± 0.25
65.95 ± 0.25
65.95 ± 0.25
65.95 ± 0.25
Early GCL
67.53 ± 0.21
66.66 ± 0.28
66.59 ± 0.29
66.63 ± 0.30
66.42 ± 0.28
Mid GCL
67.91 ± 0.19
67.04 ± 0.24
66.97 ± 0.24
67.01 ± 0.36
66.80 ± 0.29
Late GCL
68.32 ± 0.20
67.45 ± 0.23
67.38 ± 0.29
67.42 ± 0.29
67.21 ± 0.31
Early + Mid
67.62 ± 0.23
66.75 ± 0.24
66.68 ± 0.21
66.72 ± 0.27
66.51 ± 0.29
Mid + Late
68.26 ± 0.18
67.39 ± 0.23
67.32 ± 0.23
67.36 ± 0.31
67.15 ± 0.27
Appendix
Table 6 : Performance comparison of different GCL integration strategies with various kernels.
Model
Silhouette
SepRatio
Confidence
Baseline
+ GCR
Baseline
+ GCR
Baseline
+ GCR
DenseNet-121
0.4724
0.5001
2.2278
2.3325
0.9746
0.9805
ShuffleNet
0.2806
0.4083
1.7692
2.0472
0.9568
0.9619
SqueezeNet
-0.1245
-0.0825
1.0008
1.0494
0.9603
0.9660
ResNet-34
0.6032
0.7314
3.1015
4.4144
0.9801
0.9870
ResNet-50
0.5314
0.6186
2.5480
3.2294
0.9789
0.9835
Appendix
Table 7 : Quantitative metrics showing improvements in feature clustering and confidence with GCR.
plane
auto
bird
cat
deer
dog
frog
horse
ship
truck
plane
0.93
0.01
0.02
0.01
0.03
0.01
auto
0.01
0.96
0.03
bird
0.02
0.88
0.03
0.02
0.01
0.02
0.01
cat
0.01
0.02
0.83
0.02
0.08
0.01
0.01
deer
0.01
0.01
0.02
0.93
0.01
0.01
0.01
dog
0.01
0.09
0.02
0.86
0.01
Appendix
Table 8 : CIFAR-10 confusion matrix for the baseline model.
plane
auto
bird
cat
deer
dog
frog
horse
ship
truck
plane
0.94
0.02
0.01
0.01
0.01
auto
0.97
0.02
bird
0.01
0.89
0.03
0.02
0.02
0.01
0.01
cat
0.01
0.01
0.85
0.01
0.08
0.01
0.01
0.01
deer
0.01
0.01
0.02
0.94
0.01
0.01
0.01
dog
0.01
0.07
0.01
0.88
0.01
Appendix
Table 9 : CIFAR-10 confusion matrix for the model with GCR .
Model
Pre-freeze
Post-freeze
Performance Drop
Model A (Early-GCR)
66.8%
66.1%
0.7%
Model B (Late-GCR)
66.4%
65.1%
1.2%
Appendix
Table 10 : Impact of early vs . late GCR on feature robustness in ShuffleNet. Pre-freeze and post-freeze top-1 accuracy are reported, along with the performance drop.
Method
CIFAR-10
CIFAR-100
ImageNet-1K
ResNet-18
CNN2GNN [ 101 ]
95.51 ± 0.42
74.80 ± 0.81
60.12 ± 1.02
CNN2GNN + GCR
95.87 ± 0.31
76.23 ± 0.38
62.47 ± 0.47
CNN2Transformer [ 101 ]
95.79 ± 0.24
77.39 ± 0.20
71.12 ± 0.35
CNN2Transformer + GCR
95.96 ± 0.35
78.23 ± 0.30
72.33 ± 0.31
ResNet-34
Appendix
Table 11 : Performance comparison on CIFAR-10, CIFAR-100, and ImageNet-1K.
Model
ViG-Ti [ 37 ]
ViG-Ti + GCR
ViG-S [ 37 ]
ViG-S + GCR
ViG-B [ 37 ]
ViG-B + GCR
Accuracy
73.9
74.9
80.4
81.7
82.3
84.0
Appendix
Table 12 : Results on ImageNet-1K.
MobileNet
ShuffleNet
SqueezeNet
Baseline
65.95 ± 0.25
70.11 ± 0.30
69.43 ± 0.27
GCR (one-hot labels)
67.35 ± 0.24
71.28 ± 0.26
70.49 ± 0.25
GCR (soft predictions, ours )
68.32 ± 0.20
71.96 ± 0.27
71.03 ± 0.24
Appendix
Table 13 : Comparison of GCR with one-hot labels vs . soft predictions on CIFAR-100 (average over 10 runs).
λ
0
0.1
0.3
0.5
0.7
1
3
5
7
10
Top-1 Acc (%)
76.76
76.80
76.87
77.32
77.61
78.38
76.74
76.02
75.90
75.37
Appendix
Table 14 : Impact of bottleneck strength λ on CIFAR-100 (ResNet-34).
τ
0
1e−5
1e−4
1e−3
1e−2
Accuracy
78.38
78.32
78.27
78.13
76.04
Appendix
Table 15 : Effect of varying τ on CIFAR-100 with ResNet-34.
τ (deg)
Acc
1st
2nd
3rd
4th
5th
6th
intra
inter
intra
inter
intra
inter
intra
inter
intra
inter
intra
inter
baseline
65.95
38.71
13.81
22.30
7.56
12.00
5.27
5.10
4.24
17.71
19.49
8.64
9.69
ours (0.00)
68.14
31.21
12.02
18.18
7.61
9.98
4.28
3.94
3.25
16.82
18.69
8.30
9.32
0.08
68.21
31.25
12.56
18.28
6.54
9.92
4.46
4.10
3.28
18.77
19.52
9.23
9.74
0.26
67.68
36.28
13.11
20.93
7.15
10.94
4.92
4.58
3.81
17.72
19.08
8.72
9.50
0.81
67.37
35.39
13.72
20.44
6.92
11.02
4.89
4.70
3.89
17.74
18.87
8.70
9.41
Appendix
Table 16 : Results on CIFAR-100 (100 classes) with MobileNet. Accuracy and intra-/inter-class variance across six layers under different τ .
τ (deg)
Acc
1st
2nd
3rd
4th
5th
6th
intra
inter
intra
inter
intra
inter
intra
inter
intra
inter
intra
inter
baseline
77.46
30.30
8.86
17.94
4.30
9.25
2.97
3.01
2.46
9.09
12.27
4.45
6.12
ours (0.00)
77.54
30.36
8.65
17.95
4.43
9.34
2.99
3.22
2.56
9.69
12.57
4.73
6.27
0.08
77.72
31.59
9.22
18.24
4.42
9.57
3.08
3.33
2.61
9.91
12.56
4.87
6.27
0.26
77.36
32.04
9.80
18.67
4.58
9.57
3.07
3.28
2.70
9.85
12.69
4.83
6.33
0.81
77.69
31.36
9.59
18.17
4.38
9.50
3.04
3.15
2.60
9.70
12.63
4.77
6.30
Appendix
Table 17 : Results on CIFAR-100 (20 super-classes) with MobileNet. Accuracy and intra-/inter-class variance across six layers under different τ .
Conditional generative models, particularly diffusion-based methods, have recently been applied to graph prediction by modeling the target as a conditional distribution given the input graph, yielding competitive results compared to deterministic predictor. However, existing diffusion-based prediction methods typically require expensive iterative denoising at inference and often suffer from unstable sampling, which motivates recent efforts to reduce inference denoising steps and enable stable sampling via techniques such as consistency training. Despite this progress, we find that existing consistency training methods for graph prediction could potentially fall into a shortcut solution: the model may attempt to satisfy the self-consistency constraint by ignoring the noisy target (i.e., assigning it negligible weight), ultimately collapsing into a purely deterministic predictor. To mitigate such shortcut solution, we propose GCCM, a graph contrastive consistency model that goes beyond isolated pairwise matching between the same target at different noise levels by introducing negative pairs into a contrastive consistency objective. This adds an additional separation requirement, making the shortcut solution no longer trivially sufficient to satisfy the proposed objective. Moreover, we apply feature perturbation to the input node/edge features to break identical conditioning on the input graph, so that the shortcut no longer yields the same predictions across noise levels and becomes less attractive. Extensive experiments on benchmark datasets demonstrate that GCCM mitigates the shortcut solution and yields consistent performance improvements in graph prediction compared to deterministic predictors.
Shaozhen Ma, Wei Huang, Hanchen Wang +2
University of New South Wales · University of Technology Sydney
We introduce a constrained two-view framework for node prediction that aligns structure-conditioned GNN embeddings with a structure-free feature prior learned by an anchor model. Conventional Graph Neural Networks (GNNs) couple feature transformation and neighborhood aggregation, which renders them vulnerable to topology noise and heterophilous connections. To decouple this dependency, our framework utilizes an independent anchor network to capture intrinsic attribute features via a self-supervised reconstruction objective. Furthermore, we propose a Channel-Split Adaptive Gated GNN (CSAG-GNN) that dynamically routes representations between global spectral smoothing and local spatial discrimination through a node-wise gating mechanism. We propose a stable cyclic alternating optimization strategy to solve the resulting coupled bi-level objective, preventing mutual representation drift during training. Empirical results on both homophilous and heterophilous benchmarks show balanced performance gains and structural robustness over competitive baselines.
Chengcheng Yan, Qingsong Wang
College of Artificial Intelligence, Shaoxing Institute of Technology, Shaoxing, 312000, China. · School of Mathematics and Computational Science, Xiangtan University, Xiangtan, 411105, China.
Self-supervised Continual Graph Learning (CGL) aims to successively learn from a graph sequence with different tasks without label supervision - a paradigm that has attracted widespread attention. Most existing self-supervised CGL methods rely on instance-level consistency objectives that enforce stability of individual node (or node-pair) embeddings. Due to optimizing nodes in isolation, these methods fail to maintain global relational structure, causing inter-node correspondences to progressively distort under continual learning. To this end, we propose a novel Structure-Aware Optimal Transport (SAOT) framework that explicitly captures and preserves relational structure within graph representations across sequential tasks. Specifically, SAOT leverages optimal transport theory to capture global inter-node correspondences, thereby facilitating and enhancing graph representation learning. Simultaneously, SAOT incorporates a cross-task knowledge distillation mechanism to preserve the previous structural knowledge. Extensive experiments on four CGL benchmark datasets demonstrate that SAOT outperforms existing self-supervised baselines. In particular, SAOT achieves significant performance gains, improving average accuracy by up to 5% on CoraFull-CL and over 15% on Products-CL compared with state-of-the-art methods in the Class-IL setting.
Yuting Zhang, Yanbei Liu, Zhitao Xiao +3
School of Electronics and Information Engineering, Tiangong University, Tianjin, China · School of Life Sciences, Tiangong University, Tianjin, China · School of Electrical and Infomation Engineering, Tianjin University, Tianjin, China +1