cs.CLOct 8, 2026

SciTBERT: A family of chronologically consistent language models for scientific and technological language processing

Authors: Thomas Gebhart, Russell J. Funk

Organizations: Carlson School of Management, University of Minnesota

Abstract

Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

    May 22, 2026Amirhossein Yousefiramandi, Ciaran CooneyLegal IRClustering

  2. Citation-Driven Multi-View Training for Patent Embeddings: QaECTER and Sophia-Bench

    Apr 24, 2026Younes Djemmal, You Zuo, Kim Gerdes +1Legal IRText Embedding Models

  3. TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish

    Dec 28, 2025Melikşah Türker, A. Ebrar Kızıloğlu, Onur Güngör +1Language Model PretrainingLong-Context Language Modeling