GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation
Organizations: University of Houston, Texas,USA · BASIRA Lab, Department of Computing, Imperial College London, United Kingdom · Department of Mathematics and Computer Science, Faculty of Science, Alexandria University, Alexandria, Egypt · Bogazici University, Turkey · Ankara Yıldırım Beyazıt University, Turkey · ESI, Algiers, Algeria · University of Algiers 1, Algeria · York St John University London, UK · Air University, Islamabad, Pakistan · LSISI, ENSA, Mohammed Premier University, Oujda, Morocco · Universidad Católica San Pablo, Arequipa, Peru · Busitema University, Uganda · University of Laghouat, Algeria · ISI, University of Tunis El Manar, Tunisia · Tribhuvan University, Nepal · IIT Roorkee, India · Shobhit Institute of Engineering and Technology, India · MIT World Peace University, Pune, India · University of Kinshasa, DR Congo · University of Bertoua, Cameroon · University of the Witwatersrand, Johannesburg, South Africa · African Institute for Mathematical Sciences, Research and Innovation Centre (AIMS RIC), Kigali, Rwanda · KNUST, Ghana · NSUT Delhi, India · ISIMM Monastir, Tunisia · Addis Ababa University, Ethiopia
Abstract
Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.
Figures & tables
| # | Competition | Task Description | Cat. | Diff. | N | E | Feat. | Domain | Metric | Dataset |
| C01 | PROVEN-GNN | Graph classification on source code representations | Hetero | Med | 260K | 1.7M | 527 | Cybersecurity | Macro-F1 | DiverseVul |
| C02 | CGCC | Graph classification of city street-network layouts | Homo | Hard | 9.396K | 13.533K | 3 | Urban | Macro-F1 | OpenStreetMap |
| C03 | Real Or Fake! | Binary graph classification for fake news detection | Hier | Easy | 314K | 308K | 1078 | Social | Binary F1 | GossipCop |
| C04 | GLIMPS-GNN | Inductive cfRNA preeclampsia classification | Gen | Hard | 320 | 3200 | 6650 | Biomedical | F1 | GEO |
| C05 | GNN_BACE | Molecular graph classification for BACE-1 inhibition | Mol | Med | 32 | 34 | 8 | Medical | Macro-F1 | BACE (MoleculeNet) |
| C06 | DiaGraph | Binary node classification for diabetes prediction | Homo | Hard | 96K | 1.2M | 12 | Medical | Macro-F1 | Kaggle Diabetes |
| # | Competition | Diff. | P | AVG | STD | High | Low | Metric |
| C01 | PROVEN-GNN | Med | 17 | 0.792 | 0.101 | 0.900 | 0.435 | F1 |
| C02 | CGCC | Hard | 17 | 0.383 | 0.086 | 0.537 | 0.205 | F1 |
| C03 | Real Or Fake! | Easy | 15 | 0.945 | 0.053 | 0.980 | 0.767 | F1 |
| C04 | GLIMPS-GNN | Hard | 16 | 0.389 | 0.094 | 0.600 | 0.28 | F1 |
| C05 | GNN_BACE | Med | 14 | 0.496 | 0.088 | 0.674 | 0.391 | F1 |
| C06 | DiaGraph | Hard | 15 | 0.570 | 0.125 | 0.822 | 0.477 | F1 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| # | Model Name | Model Engine Name | Source |
| Proprietary Chat LLMs | |||
| 1 | Claude Opus 4.6 Anthropic [2026a] | claude-opus-4-6 | Link |
| 2 | Gemini-3 Flash Google DeepMind [2026] | gemini-3.0-flash | Link |
| Open-source Chat LLMs | |||
| 3 | Llama-3.3 70B Meta AI [2024] | Llama-3.3-70B-Instruct | Link |
| 4 | Qwen2.5-Coder 32B Alibaba Cloud [2024] | Qwen2.5-Coder-32B-Instruct | Link |