SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
Organizations: Carlson School of Management, University of Minnesota
Abstract
Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.
Figures & tables
| Format | Name | Train | Test | Metric | Source |
| CLF | CPC Section | 1,041,520 | 260,369 | Macro F1 | USPTO |
| CPC Subclass | 683,960 | 170,978 | Macro F1 | USPTO | |
| Assignee Type | 940,533 | 235,121 | Macro F1 | USPTO | |
| Academic Public | 182,955 | 45,670 | Binary F1 | OCPB | |
| US Foreign | 1,777,085 | 444,143 | Binary F1 | OCPB | |
| VC Backed | 804,792 | 201,106 | Binary F1 | OCPB |
| Classification (F1) | Regression ( ) | Proximity (MAP) | |||||||
| Model | Sci | Pat | Mean | Sci | Pat | Mean | Sci | Pat | Mean |
| SPECTER2 | |||||||||
| SciBERT | |||||||||
| SciNCL | |||||||||
| PaECTER | |||||||||
| ChronoBERT-2013 | |||||||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Nodes (M) | Edges (M) | Cross-domain pos. (%) | ||||||
|---|---|---|---|---|---|---|---|---|
| Cutoff | Papers | Patents | S S | P P | P S | Anchors | All | P S pool |
| Setting | Value |
|---|---|
| node2vec | |
| Embedding dimension | |
| Walk length / context size | / |
| Walks per node | |
| Negative samples | |
| Return / in-out ( / ) | / |
| Cutoff | S2ORC | Abstracts | USPTO | FineWeb-Edu | Total |
|---|---|---|---|---|---|
| Cutoff | Warmup | Stable | Decay | Budget | Steps |
|---|---|---|---|---|---|
| Cutoff | S2ORC (47.1%) | Abstracts (11.8%) | USPTO (35.3%) | FineWeb-Edu (5.9%) |
|---|---|---|---|---|
| ( ) | ( ) | ( ) | ( ) | |
| ( ) | ( ) | ( ) | ( ) | |
| ( ) | ( ) | ( ) | ( ) | |
| ( ) | ( ) | ( ) | ( ) | |
| ( ) | ( ) | ( ) | ( ) | |
| ( ) | ( ) | ( ) | ( ) |
| Source | Used for | License |
|---|---|---|
| USPTO (via PatentsView) | patent text; CPC, assignee, inventor, citation, and maintenance fee labels | CC BY 4.0 |
| Kogan et al. [2017] | Kogan Value | None stated |
| Reliance on Science [ Marx and Fuegi, 2020 ] | patent-to-paper citations | CC BY-NC 4.0 |
| Patent–paper pairs [ Marx and Fuegi, 2020 ] | PPP tasks | CC BY-NC 4.0 |
| Pasteur’s quadrant researchers [ Scharfmann et al., 2025 ] | PQR tasks | CC BY-NC 4.0 |
| Assignee attributes [ Ewens and Marx, 2024 ] | ownership tasks, Same Initial Assignee | CC BY-NC 4.0 |
| Classification (F1) | Regression (Kendall ) | Proximity (MAP) | Proximity (NDCG) | |||||||||||||||
| Model | Biomimicry | DRSM | Fields of study | Citation Count | Max hIndex | Peer Review Score | Publication Year | Tweet Mentions | Highly Influential Citations | Same Author Detection | SciDocs Cite | SciDocs CoCite | SciDocs CoRead | SciDocs CoView | NFCorpus | RELISH | Search | TREC-CoVID |
| SPECTER2 | ||||||||||||||||||
| SciBERT | ||||||||||||||||||
| SciNCL | ||||||||||||||||||
| PaECTER | ||||||||||||||||||
| ChronoBERT-2013 | ||||||||||||||||||
| Model | Academic Public | Assignee Type | CPC Section | CPC Subclass | Grant Abandon | PPP Paper | PPP Patent | PQR Paper | PQR Patent | Renewal | US Foreign | VC Backed |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SPECTER2 | ||||||||||||
| SciBERT | ||||||||||||
| SciNCL | ||||||||||||
| PaECTER | ||||||||||||
| ChronoBERT-2013 | ||||||||||||
| ChronoBERT-2024 |
| Regression ( ) | Proximity (MAP) | ||||||||||||||
| Model | Forward Citation | Grant Year | Kogan Value | Science Uptake | Assignee Patent Match | Paper Patent Retrieval | Papers Cocite | Patent Cite | Patent Cocite | Patent Paper Citation | PPP Paper Patent | PQR Paper Patent | PQR Patent Paper | Same Initial Assignee | Same Inventor |
| SPECTER2 | |||||||||||||||
| SciBERT | |||||||||||||||
| SciNCL | |||||||||||||||
| PaECTER | |||||||||||||||
| ChronoBERT-2013 | |||||||||||||||