cs.LGOct 4, 2026

GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation

Authors: Murad Hossen, Tasneem Selim, Gurur Gamgam, Tuga Yousif, Abderrahmane Kasmi, Ikram Aissiou, Mubaraq Onipede, Faran Taimoor Butt, +17 more

Organizations: University of Houston, Texas,USA · BASIRA Lab, Department of Computing, Imperial College London, United Kingdom · Department of Mathematics and Computer Science, Faculty of Science, Alexandria University, Alexandria, Egypt · Bogazici University, Turkey · Ankara Yıldırım Beyazıt University, Turkey · ESI, Algiers, Algeria · University of Algiers 1, Algeria · York St John University London, UK · Air University, Islamabad, Pakistan · LSISI, ENSA, Mohammed Premier University, Oujda, Morocco · Universidad Católica San Pablo, Arequipa, Peru · Busitema University, Uganda · University of Laghouat, Algeria · ISI, University of Tunis El Manar, Tunisia · Tribhuvan University, Nepal · IIT Roorkee, India · Shobhit Institute of Engineering and Technology, India · MIT World Peace University, Pune, India · University of Kinshasa, DR Congo · University of Bertoua, Cameroon · University of the Witwatersrand, Johannesburg, South Africa · African Institute for Mathematical Sciences, Research and Innovation Centre (AIMS RIC), Kigali, Rwanda · KNUST, Ghana · NSUT Delhi, India · ISIMM Monastir, Tunisia · Addis Ababa University, Ethiopia

Abstract

Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

    Aug 2, 2026Fali Wang, Ali Al-Lawati, Iliyas Bektas +5Reasoning BenchmarkMulti-Turn Benchmark

  2. GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs

    Sep 14, 2026Zixiang Xu, Yanbo Wang, Chenxi Wang +6Graph RepresentationsLarge Language Model Benchmarks

  3. GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks

    Aug 3, 2026Jiarui Tan, Zhongjian Zhang, YaBo Guo +5Large Language Model Agents