Apr 16, 2026, cs.CLJ/K move · Enter open · S save
Rami Luisto, Liisa Petäinen, Tommi Grönholm, Jan Böhm+4
Faculty of Information Technology, University of Jyväskylä, Jyväskylä, Finland. · Digital Workforce Services, Helsinki, Finland. · Heart and Lung Center, Helsinki University Hospital, Helsinki, Finland.+2
In Natural Language Processing (NLP) classification tasks where a lack of labeled data is an issue, continued pretraining (CPT) of transformer models on unlabeled data is an established approach. In this paper, we have two aims. (1) We describe our observations from continued pretraining of the Finnish BERT transformer model (FinBERT) on a Finnish histopathological dataset (below, \emph{the Histopathology data}). (2) Since the Histopathology data has no classification labels, we gather public Finnish datasets as proxy data to analyze whether the signals observed in (1) are associated with downstream classification gains. We observe that CPT train-time loss curves differ strongly by domain, and that, in an exploratory analysis, certain CPT-derived features correlate with proxy classification improvement. In particular, this report contributes to the limited literature on NLP for Finnish healthcare data.