Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.
Figures & tables
Figure 1: Performance of SciTBERT and SciTBERT-CI models, alongside a number of prior encoder models, across both SciRepEval and PatRepEval benchmarks. Performance is aggregated by taking the mean over the tasks of each category within each benchmark, then the mean of the two benchmarks. See Table 2 .
Format
Name
Train
Test
Metric
Source
CLF
CPC Section
1,041,520
260,369
Macro F1
USPTO
CPC Subclass
683,960
170,978
Macro F1
USPTO
Assignee Type
940,533
235,121
Macro F1
USPTO
Academic Public
182,955
45,670
Binary F1
OCPB
US Foreign
1,777,085
444,143
Binary F1
OCPB
VC Backed
804,792
201,106
Binary F1
OCPB
Table 1: Summary of PatRepEval tasks across the three formats: classification (CLF), regression (RGN), and proximity (PRX). Proximity tasks are evaluated by retrieval over a corpus, so they have Q queries ranked against a corpus of C documents and no training split. Tasks marked † take scientific papers as input; all other tasks take patents. Label sources are USPTO: U.S. Patent and Trademark Office; KPSS: Kogan et al. (2017) ; RoS: Reliance on Science ( Marx and Fuegi, 2020 ) ; PPP: patent–paper pairs ( Marx and Fuegi, 2020 ) ; PQR: Pasteur’s quadrant researchers ( Scharfmann et al., 2025 ) ; OCPB: assignee attributes ( Ewens and Marx, 2024 ) .
Classification (F1)
Regression ( τ )
Proximity (MAP)
Model
Sci
Pat
Mean
Sci
Pat
Mean
Sci
Pat
Mean
SPECTER2
63.5
66.1
64.8
27.7
30.5
29.1
77.4
15.9
46.7
SciBERT
63.9
65.7
64.8
26.7
30.4
28.6
68.8
10.2
39.5
SciNCL
65.1
65.8
65.4
26.5
29.7
28.1
79.7
16.0
47.8
PaECTER
61.4
68.2
64.8
22.1
35.7
28.9
77.1
21.7
49.4
ChronoBERT-2013
62.8
65.6
64.2
26.5
29.1
27.8
67.6
8.5
38.1
Table 2: Mean performance by task category. Sci and Pat are the means over the SciRepEval and PatRepEval tasks of each category, and the Mean column averages those means. All scores are on a 0 – 100 scale. The best performance in each column is bold and the second is underlined.
Figure 2: Performance of 2013 and 2024 vintage SciTBERT and SciTBERT-CI models, alongside a number of prior encoder models, across a selection of PatRepEval tasks.
Figure 3: Performance of odd-year SciTBERT and SciTBERT-CI models, alongside a number of baseline models, across a selection of SciRepEval tasks.
Figure 4: Relative performance of SciTBERT and SciTBERT-CI models versus similar scientific and chronologically timestamped models on all PatRepEval and SciRepEval tasks. Pat indicates PatRepEval tasks and Sci indicates SciRepEval tasks. Performance is reported in standard deviations from the task mean performance across the models. The first-, second-, and third-best performing model for the task is denoted with a 1, 2, or 3, respectively.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Nodes (M)
Edges (M)
Cross-domain pos. (%)
Cutoff y
Papers
Patents
S → S
P → P
P → S
Anchors
All
P → S pool
2013
2.7
4.6
27.3
49.1
0.6
348,242
1.7
3.9
2014
3.3
4.9
33.5
54.8
0.7
348,316
1.6
3.9
2015
3.9
5.1
41.0
60.1
0.8
348,534
1.6
3.8
2016
4.6
5.4
50.1
65.3
0.9
348,557
1.6
4.2
2017
5.4
5.7
61.4
71.0
1.0
348,568
1.6
4.3
Appendix
Table 3: The post-training citation graph at each year cutoff y . Nodes are papers and patents with a title and an abstract, and an edge is included only if both of its endpoints are dated no later than y . The last two columns give the share of positives whose type differs from the anchor’s, over all anchors and over anchors drawn from the pool of patents that cite at least one paper.
Setting
Value
node2vec
Embedding dimension
128
Walk length / context size
20 / 10
Walks per node
10
Negative samples
5
Return / in-out ( p / q )
1 / 1
Appendix
Table 4: Hyperparameter details for the node2vec, sampling, and contrastive learning steps of the SciTBERT-CI post-training. All settings are shared by each SciTBERT- y -CI model vintage.
Cutoff y
S2ORC
Abstracts
USPTO
FineWeb-Edu
Total
2013
10.7
0.8
10.2
23.2
44.9
2014
12.9
1.0
10.8
117.3
142.0
2015
15.5
1.2
11.4
228.0
256.1
2016
18.4
1.5
12.1
329.6
361.6
2017
21.8
1.7
12.7
504.2
540.5
2018
25.8
2.3
13.4
667.5
708.9
Appendix
Table 5: Cumulative pretraining corpus C≤y available at each year cutoff, in billions of tokens. Counts are estimated by scaling each source’s whitespace token count by a byte-pair-encoding factor calibrated on a 10,000 -document sample and by its measured language-ID retention rate. FineWeb-Edu is reported from 2013 , the first Common Crawl dump ingested, and paper and patent text from 2000 . Papers and patents predating 2000 are therefore uncounted in this table but make up a small portion of overall documents.
Cutoff y
Warmup
Stable
Decay
Budget
Steps
2013
0.5
20.0
3.0
23.5
44,821
2014
0.6
23.4
3.5
27.5
52,450
2015
0.7
27.6
4.1
32.4
61,797
2016
0.7
29.1
4.4
34.2
65,230
2017
0.8
30.7
4.6
36.1
68,853
2018
0.8
32.2
4.8
37.8
72,096
Appendix
Table 6: Per-cutoff token budgets, in billions of tokens, and the resulting optimizer steps. Each budget is set to roughly exhaust the primary scientific and patent sources under the sampling mixture, so it grows with C≤y . The phase split is fixed at 2.1% / 85.1% / 12.8% throughout.
Cutoff y
S2ORC (47.1%)
Abstracts (11.8%)
USPTO (35.3%)
FineWeb-Edu (5.9%)
2013
11.1 ( 1.03× )
2.8 ( 3.36× )
8.3 ( 0.81× )
1.4 ( 0.06× )
2014
12.9 ( 1.00× )
3.2 ( 3.23× )
9.7 ( 0.90× )
1.6 ( 0.01× )
2015
15.2 ( 0.98× )
3.8 ( 3.15× )
11.4 ( 1.00× )
1.9 ( <0.01× )
2016
16.1 ( 0.87× )
4.0 ( 2.77× )
12.1 ( 1.00× )
2.0 ( <0.01× )
2017
17.0 ( 0.78× )
4.2 ( 2.45× )
12.7 ( 1.00× )
2.1 ( <0.01× )
2018
17.8 ( 0.69× )
4.4 ( 1.96× )
13.3 ( 1.00× )
2.2 ( <0.01× )
Appendix
Table 7: Tokens consumed from each source during pretraining, in billions, with the number of passes over the source in parentheses. The sampling mixture is fixed, so consumption is the token budget times the source’s sampling fraction.
Source
Used for
License
USPTO (via PatentsView)
patent text; CPC, assignee, inventor, citation, and maintenance fee labels
CC BY 4.0
Kogan et al. [2017]
Kogan Value
None stated
Reliance on Science [ Marx and Fuegi, 2020 ]
patent-to-paper citations
CC BY-NC 4.0
Patent–paper pairs [ Marx and Fuegi, 2020 ]
PPP tasks
CC BY-NC 4.0
Pasteur’s quadrant researchers [ Scharfmann et al., 2025 ]
PQR tasks
CC BY-NC 4.0
Assignee attributes [ Ewens and Marx, 2024 ]
ownership tasks, Same Initial Assignee
CC BY-NC 4.0
Appendix
Table 8: Data sources used to construct PatRepEval and the terms under which each is distributed. Most of the USPTO data is derived from PatentsView [ Toole et al., 2021 ] .
Classification (F1)
Regression (Kendall τ )
Proximity (MAP)
Proximity (NDCG)
Model
Biomimicry
DRSM
Fields of study
Citation Count
Max hIndex
Peer Review Score
Publication Year
Tweet Mentions
Highly Influential Citations
Same Author Detection
SciDocs Cite
SciDocs CoCite
SciDocs CoRead
SciDocs CoView
NFCorpus
RELISH
Search
TREC-CoVID
SPECTER2
74.3
74.4
41.7
37.6
15.4
21.7
36.4
27.4
42.4
84.8
83.2
88.0
83.7
82.4
62.7
88.9
73.2
87.3
SciBERT
73.0
74.5
44.0
37.1
16.1
22.3
31.3
26.8
39.5
82.7
70.4
76.0
70.8
73.3
56.0
85.6
71.3
82.8
SciNCL
74.5
75.8
44.9
37.1
14.3
22.8
31.2
27.2
42.7
86.1
90.0
89.9
85.9
83.6
69.2
89.1
73.4
86.6
PaECTER
71.4
72.8
40.1
30.9
11.2
17.8
25.4
25.1
43.4
82.7
84.2
87.0
82.8
82.8
63.4
89.8
72.1
86.9
ChronoBERT-2013
69.8
74.4
44.1
37.1
15.5
24.4
28.7
26.6
39.4
82.1
69.3
73.2
70.0
71.6
54.8
85.1
71.3
84.1
Appendix
Table 9: SciRepEval performance, per task. All scores are on a 0 – 100 scale. Bold denotes best-performing model for the task and underline the second (columnwise).
Model
Academic Public
Assignee Type
CPC Section
CPC Subclass
Grant Abandon
PPP Paper
PPP Patent
PQR Paper
PQR Patent
Renewal
US Foreign
VC Backed
SPECTER2
79.8
25.8
67.3
59.5
63.8
66.6
86.8
64.6
75.5
60.5
75.9
66.9
SciBERT
79.6
26.3
64.7
56.7
63.9
66.6
86.4
65.0
75.1
61.4
76.1
66.3
SciNCL
79.8
26.2
68.4
60.0
63.0
63.7
86.6
64.9
73.9
60.3
75.6
66.9
PaECTER
81.9
28.9
77.1
68.1
64.9
61.9
86.8
62.7
75.6
62.9
79.5
68.3
ChronoBERT-2013
79.0
26.2
66.1
57.0
64.0
66.3
86.0
64.5
75.0
61.1
76.4
65.8
ChronoBERT-2024
79.1
26.3
66.4
57.0
63.9
66.2
85.9
64.7
75.5
61.8
75.3
65.9
Appendix
Table 10: PatRepEval classification performance, per task. All scores are on a 0 – 100 scale. Bold denotes best-performing model for the task and underline the second (columnwise).
Regression ( τ )
Proximity (MAP)
Model
Forward Citation
Grant Year
Kogan Value
Science Uptake
Assignee Patent Match
Paper Patent Retrieval
Papers Cocite
Patent Cite
Patent Cocite
Patent Paper Citation
PPP Paper Patent
PQR Paper Patent
PQR Patent Paper
Same Initial Assignee
Same Inventor
SPECTER2
24.6
38.5
38.7
20.1
17.8
8.9
8.1
10.0
11.0
3.8
74.6
2.7
2.0
25.5
10.4
SciBERT
24.9
38.5
38.6
19.7
13.9
2.9
4.5
6.6
7.7
1.4
43.2
1.1
0.8
21.8
8.1
SciNCL
24.1
36.7
38.4
19.5
16.8
9.3
8.4
10.7
11.2
4.6
75.0
2.9
1.9
25.1
10.1
PaECTER
30.8
49.1
43.8
19.0
23.3
14.3
9.9
21.0
20.8
8.1
87.3
3.9
4.4
30.9
15.0
ChronoBERT-2013
23.9
35.7
38.2
18.7
11.9
1.7
3.7
6.3
7.3
1.0
33.1
0.6
0.5
20.5
7.4
Appendix
Table 11: PatRepEval regression and proximity performance, per task. All scores are on a 0 – 100 scale. Bold denotes best-performing model for the task and underline the second (columnwise).
Figure 5: Relative performance of SciTBERT models versus baseline MLM-only scientific and timestamped models across both benchmark tasks. Performance is reported in standard deviations from the task mean performance across the models. The first-, second-, and third-best performing model for the task is denoted with a 1, 2, or 3, respectively.
Figure 6: Relative performance of SciTBERT-CI models versus similar contrastively post-trained scientific models. Performance is reported in standard deviations from the task mean performance across the models. The first-, second-, and third-best performing model for the task is denoted with a 1, 2, or 3, respectively.