cs.LGOct 6, 2025

Fractional Heat Kernel for Semi-Supervised Graph Learning with Small Training Sample Size

Authors: Farid Bozorgnia, Vyacheslav Kungurtsev, Shirali Kadyrov, Mohsen Yousefnezhad

Abstract

We develop a source-driven fractional heat-kernel framework for semi-super-vised graph learning that combines nonlocal propagation with sustained label information. A fixed nonzero label source compatible with the Laplacian null space prevents asymptotic collapse into that space, providing a mechanism for mitigating oversmoothing at long diffusion times. The fractional order controls the relative modal attenuation and the spectral weighting of the sustained response, while the diffusion time sets the propagation horizon. We characterize conservation laws and equilibria on normalized and disconnected graphs, develop a null-space deflation, and analyze the approximation of the propagators. On Two-Moon, fractional orders improve source-free propagation, while compatible source-driven diffusion exceeds 96%96\% mean accuracy with one training label per class at orders 0.80.8 and 11. On Cora and CiteSeer, the source-driven pipeline improves mean accuracy over GAT by 9.29.2 and 8.08.0 percentage points at one training label per class, with model selection on 500500 labeled validation nodes and closely comparable classical and fractional pipeline configurations; it matches GAT on PubMed and trails it at ten and twenty labels per class on Cora. Within GraphHeat, validation-based exponent selection at a common diffusion time yields a paired gain of 1.451.45 percentage points at one label per class.

Figures & tables

Explore similar work

Apr 22, 2026cs.LG

F\textsuperscript{2}LP-AP: Fast & Flexible Label Propagation with Adaptive Propagation Kernel

Semi-supervised node classification is a foundational task in graph machine learning, yet state-of-the-art Graph Neural Networks (GNNs) are hindered by significant computational overhead and reliance on strong homophily assumptions. Traditional GNNs require expensive iterative training and multi-layer message passing, while existing training-free methods, such as Label Propagation, lack adaptability to heterophilo-us graph structures. This paper presents \textbf{F2^2LP-AP} (Fast and Flexible Label Propagation with Adaptive Propagation Kernel), a training-free, computationally efficient framework that adapts to local graph topology. Our method constructs robust class prototypes via the geometric median and dynamically adjusts propagation parameters based on the Local Clustering Coefficient (LCC), enabling effective modeling of both homophilous and heterophilous graphs without gradient-based training. Extensive experiments across diverse benchmark datasets demonstrate that \textbf{F2^2LP-AP} achieves competitive or superior accuracy compared to trained GNNs, while significantly outperforming existing baselines in computational efficiency.
May 27, 2026cs.LG

Where LLM Annotators Fail: Label-Free Learning on Graphs with LLMs

Node classification on graphs often requires labeled nodes, yet obtaining labels at graph scale is expensive. When node attributes contain semantic content, such as paper abstracts, web pages, or product descriptions, large language models (LLMs) can provide low-cost supervision by annotating a small subset of nodes. However, these LLM-generated labels are noisy, and existing label-free graph learning methods usually treat this noise as either global or class-conditional. We find that LLM annotation errors are not only class-dependent but also region-dependent: within the same class, reliability can vary sharply across feature-space clusters. In light of this, we propose Cluster-Aware Noise Estimation (CANE), a label-free learning framework that estimates cluster-conditional LLM reliability without ground truth labels, and uses this estimate to decide which pseudo-labels to trust, and which labels to correct. Across various graph benchmarks and GNN backbones, CANE improves over the strongest label-free baselines, with the largest gains on datasets exhibiting stronger cluster-conditional noise.
Sep 28, 2026cs.LG

Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers

We study post-hoc refinement of frozen node classifiers: given only the graph GG and class distributions QQ predicted by a frozen model, can we improve accuracy without access to node features, model parameters, or gradients? APPNP answers this by propagating logits with a restart towards the initial predictions, minimizing the anchored Dirichlet energy. Instead, we consider the Potts energy, and decompose it into a Dirichlet term, which penalizes disagreement between neighbouring nodes, and a Gini term, which penalizes indecision within each node. This decomposition motivates Propagate, Then Sharpen (PtS), which alternates between propagation of class probabilities and node-wise, mass-preserving sharpening, with only one additional hyperparameter selected using labelled validation nodes. Across nine homophilic graphs, with a frozen MLP backbone, PtS improves mean test accuracy over independently tuned APPNP by 1.711.71 percentage points on clean inputs and 3.903.90 under severe Gaussian feature corruption. Gains over APPNP become smaller, but remain positive with frozen GCN and GraphSAGE backbones. Sharpening also removes most of the accuracy loss of deep propagation: on clean inputs without restart, accuracy falls by 2.22.2 points between 22 and 100100 propagation steps under PtS, compared with 33.833.8 for APPNP.