cs.CVAug 3, 2026

Modeling Scientific Experiment Scenes: Dataset and Model

Authors: Minghao ZouQingtian ZengShangkun LiuCong LiuPaul L. RosinGuanghui YueJun LiuWei Zhou

Organizations: College of Computer Science and Engineering, Shandong University of Science and Technology, Qingdao, China · School of Computer Science and Informatics, Cardiff University, Cardiff, UK · ABC FINTECH Company Limited · NOVA Information Management School, Universidade Nova de Lisboa, Lisboa, Portugal · School of Biomedical Engineering, Shenzhen University, China · School of Computing and Communications, Lancaster University, UK

Abstract

Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily-life images and overlook scientific experiment scenes with specialized instruments, task-specific experimental semantics, and dense, fine-grained physical relations. Building upon PhysScene, our previously introduced SGG dataset for physics experiment scenes, we further identify two key challenges that such scientific environments pose to existing SGG models: a pronounced long-tail relational predicate distribution and a substantial visual-textual semantic gap. To address these challenges, we propose the Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG. The model enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues. We also incorporate relation-aware pre-training, caption-derived pseudo-supervision, and adaptive weighting to support balanced learning across head and tail predicates. Extensive experiments on PhysScene and VG150 show that CM-DPG achieves competitive performance across multiple evaluation settings, with ablation studies validating the contribution of each component. The dataset and code are publicly available at https://github.com/ZMH-SDUST/CM-DPG.

Explore similar work

CardsList