SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
Authors: Jennifer D'Souza, Sameer Sadruddin, Anisa Rula, Ana Bossler, Andrés Fullana, Enric Bas, Syed Ather, Defne Circi, +18 more
Organizations: TIB Leibniz Information Centre for Science and Technology, Hannover, Germany · University of Brescia, Brescia, Italy · University of Alicante, Alicante, Spain · Georgia Institute of Technology, Atlanta, United States · Duke University, Durham, United States · Johns Hopkins University, Baltimore, United States · University of Manchester, Manchester, United Kingdom · Narnia Labs, Daejeon, South Korea · Hewlett Packard Enterprise Labs, India · Wismar University of Applied Sciences, Wismar, Germany · University of Rostock, Rostock, Germany · University College London, London, United Kingdom · University of Milano-Bicocca, Milan, Italy · Cambridge Institute of Technology, Bengaluru, India · Dangote Fertiliser Limited, Lagos, Nigeria · North Dakota State University, Fargo, United States · University of Alberta, Edmonton, Canada · SES AI, Woburn, United States
Abstract
Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of SciSchema.org, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology & Biotechnology, Materials & Chemistry, Imaging & Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.
Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually designing an annotation schema and exhaustively labeling the corpus, a slow and error-prone process. We introduce ScheMatiQ, which leverages calls to a backbone LLM to take a question and a corpus to produce a schema and a grounded database, with a web interface that lets steer and revise the extraction. In collaboration with domain experts, we show that ScheMatiQ yields outputs that support real-world analysis in law and computational biology. We release ScheMatiQ as open source with a public web interface, and invite experts across disciplines to use it with their own data. All resources, including the website, source code, and demonstration video, are available at: www.ScheMatiQ-ai.com
Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical data. Even if it is a quite recent proposal, it has been applied in several biomedical contexts. Developing and maintaining LinkML schemas presents several challenges, particularly for novice curators. Non-expert bio-curators may struggle with LinkML syntax and best practices, requiring significant time and effort to develop well-structured schemas. Results: In this paper we propose SchemaLink, a web-based environment for the graphical construction and enhancement of LinkML schemas that address the following requirements: (i) introduce a graphical language for the specification of LinkML schemas, (ii) make uniform the specification of schemas in similar contexts, (iii) simplify the design and curation processes by exploiting a RAG-based approach to assist curators in creating new schemas from scratch and editing already developed ones. Several experimental analyses show the quality of the produced LinkML schemas through the AI-based editing facilities. Availability and Implementation: SchemaLink is available online at: https://SchemaLink.biodata.di.unimi.it. SchemaLink code and testing data are available as open-source on GitHub at: https://github.com/AnacletoLAB/{schemalink-webapp,schemalink-api}.
Emanuele Cavalleri, Paolo Perlasca, J. Harry Caufield +3
Atomic layer deposition (ALD) and atomic layer etching (ALE) are reported heterogeneously across experimental and simulation literature in materials science, hindering comparison and machine-actionable reuse. We present four domain-expert-reviewed JSON Schemas for ALD and ALE experimental and simulation processes. Curated with schema-miner and grounded in QUDT using schema-miner pro, the schemas structure materials, process conditions, configurations , and measured or predicted results. We compare their scope, structure, and semantic grounding, and demonstrate their use for schema-guided literature extraction and publication of structured records through ORKG templates.