Template Ageing and Longitudinal Verification in Fixed-Text Keystroke Dynamics: A Subject-Disjoint Study Across Eight Weeks
Authors: Simon Parkinson, Saad Khan, Na Liu, Qing Xu
Organizations: Department of Computer Science, University of Huddersfield, HD1 3DH Huddersfield, U.K. · College of Intelligence and Computing, Tianjin University, Tianjin, China
Behavioural biometric templates are widely believed to degrade as the gap between enrolment and verification grows, but few studies measure this template ageing effect directly under controlled conditions. We collected a longitudinal dataset of 40 fixed passwords, each typed four times per weekly session over eight consecutive weeks. We compare a scaled-Manhattan matcher (M1), a gradient-boosted classifier (M2), a TypeNet-style recurrent embedding model (M3), and a TypeFormer-style Transformer (M4) under a 5-fold subject-disjoint protocol and a design that jointly varies mechanism and the enrolment-to-query gap, from 0 to 7 weeks. Template ageing proves large and systematic. Error increases monotonically with the gap for every mechanism, from an EER of 14.6-27.2% at a gap of zero to 25.5-37.1% at seven weeks, or 1.7% of decision error per week elapsed (p < 0.001). However, the choice of mechanism matters more than its rate of ageing. Baseline accuracy spans 12.6 percentage points across the four mechanisms, the degradation each accumulates over seven weeks spans only 2.3 points, and ageing never reorders them. A matcher can therefore be chosen on same-session accuracy, with ageing managed by re-enrolment scheduling rather than by matcher selection. The two properties are nonetheless distinct, as M3 is the least accurate mechanism yet ages significantly more slowly than M1 under every specification tested. Training randomness also matters differently by architecture, with 58% of the recurrent model's fold-to-fold variance attributable to seed noise against 19% for the Transformer. Because the smaller ageing-rate differences are sensitive to modelling choices, while the accuracy differences and the ageing effect are not, we recommend that comparative ageing-rate claims be supported by seed-level score fusion, independent replication, and an alternative outcome-model specification.
Figures & tables
Study
Longitudinal (weeks+)
Cross-week template test
Learned embedding model
Subject-disjoint evaluation
Inferential statistics
Montalvão et al. [ 27 ]
×
×
×
×
×
Acien et al. (TypeNet) [ 9 ]
× (free-text)
×
✓
✓
×
Parkinson et al. [ 12 ]
✓
×
×
×
×
Giot/Pisani (template update) [ 6 , 7 ]
∼ (multi-session)
∼ (adaptation active)
×
×
×
Yang et al. [ 8 ]
✓ (drift-detection)
×
×
×
×
This paper
✓
✓
✓
✓
✓
TABLE I: Positioning relative to selected prior work (qualitative; ✓ = present, × = absent, ∼ = partially addressed).
TABLE II: Architecture and optimisation hyperparameters for M3 (LSTM) and M4 (Transformer), fixed a priori and shared across folds and replicates.
Fig. 1: EER versus enrolment-to-query gap Δ (weeks), by mechanism and outlier-clipping condition, at k=5 enrolment sessions. Shaded bands are bootstrap 95% confidence intervals over test participants. Error increases monotonically with Δ for every mechanism.
Mechanism
EER at Δ=0
EER at Δ=7
M1 (scaled Manhattan)
21.0 [19.7, 22.4]
33.1 [30.1, 36.0]
M2 (shallow ML)
17.2 [15.9, 18.5]
27.0 [23.7, 30.5]
M3 (TypeNet-style LSTM)
27.2 [25.6, 28.8]
37.1 [34.7, 39.5]
M4 (TypeFormer-style)
14.6 [13.5, 15.9]
25.5 [21.8, 29.8]
Trials (all mechanisms): genuine / impostor / total
Δ=0
23,827 / 352,560 / 376,387
TABLE III: EER (%) at Δ=0 and Δ=7 ( k=5 , clipped, M3/M4 three-seed fusion, with 95% bootstrap CI in brackets), together with the underlying genuine/impostor trial counts (identical across mechanisms).
Fig. 2: DET curves at k=5 , pooled across Δ , for the outlier-clipped (left) and unclipped (right) conditions.
Mechanism
FNMR @ FMR=1%
FNMR @ FMR=0.1%
M1 (scaled Manhattan)
91.5
98.6
M2 (shallow ML)
79.7
93.9
M3 (TypeNet-style LSTM)
95.2
99.3
M4 (TypeFormer-style)
86.4
97.8
TABLE IV: FNMR (%) at fixed FMR operating points, k=5 , clipped, pooled across Δ .
Term
Coef.
z
p
Δ (per week, centred)
+0.017
21.60
<0.001
Mechanism: M2 vs. M1
−0.040
−18.30
<0.001
Mechanism: M3 vs. M1
+0.052
23.90
<0.001
Mechanism: M4 vs. M1
−0.065
−30.15
<0.001
Length (per character, centred)
−0.020
−11.27
<0.001
Hardware-change (week 5)
+0.005
2.08
0.037
TABLE V: Mixed-effects model: selected fixed-effect estimates (main-effects model, no interaction). N=299,984 , with participant and password variance components of 0.003 and 0.001 respectively.
Linear probability model
Logistic (cluster-robust)
Mechanism vs. M1
Run 1
Run 2
Run 3
Run 1
Run 2
Run 3
M2
−0.002 (.029)
−0.003 (.010)
−0.002 (.030)
−0.004 (.652)
−0.004 (.650)
+0.003 (.713)
M3
−0.002 (.047)
−0.004 ( < .001)
−0.004 (.001)
−0.019 (.032)
−0.021 (.016)
−0.015 (.056)
M4
−0.000 (.678)
−0.001 (.485)
−0.000 (.954)
+0.016 (.093)
+0.016 (.090)
+0.024 ( .008 )
TABLE VI: Δ× Mechanism interaction under the linear-probability and logistic (cluster-robust) specifications, across three independent end-to-end pipeline replicates. Cells show the coefficient, with the p -value in parentheses. In the linear specification, coefficients are expressed on the weekly probability-of-error scale per unit of Δ , whereas in the logistic specification, coefficients are measured in log-odds. Consequently, only the direction and statistical significance of the coefficients, rather than their magnitudes, are directly comparable across the two model formulations. For the joint test in the linear model, the test statistics are χ2(3)=7.27,17.92,and 16.22 , with corresponding p -values of 0.064,0.0005,and 0.0010 for Runs 1, 2, and 3, respectively.
Mechanism
Within-fold seed SD
Across-fold SD
Noise share
M3 (LSTM)
0.0181
0.0313
∼ 58%
M4 (Transformer)
0.0062
0.0325
∼ 19%
TABLE VII: Decomposition of fold-to-fold EER variance into training-seed noise (same data, different training run) versus total across-fold spread.
Δ
Genuine trials
Impostor trials
Total
0
23,827
352,560
376,387
1
79,994
293,714
373,708
2
68,527
249,512
318,039
3
56,387
203,999
260,386
4
45,332
163,875
209,207
5
34,393
123,530
157,923
TABLE VIII: Genuine and impostor trial counts by Δ ( k=5 , clipped; identical across mechanisms).
Mechanism
Condition
Δ=0
Δ=1
Δ=2
Δ=3
Δ=4
Δ=5
Δ=6
Δ=7
M1 (scaled Manhattan)
Clipped
21.0
24.2
25.7
27.0
28.3
29.3
30.7
33.1
Unclipped
24.0
26.3
27.7
28.8
30.0
31.2
32.2
34.3
M2 (shallow ML)
Clipped
17.2
20.6
22.1
23.5
24.5
25.2
25.6
27.0
Unclipped
19.8
23.1
24.5
25.9
27.0
27.6
27.8
28.6
M3 (TypeNet-style LSTM)
Clipped
27.2
29.0
30.5
32.1
33.4
34.0
35.2
37.1
Unclipped
21.4
24.1
25.7
27.3
28.9
30.1
31.8
34.1
TABLE IX: Full EER (%) by mechanism, clipping condition, and temporal gap Δ , k=5 .
Mechanism
k
Δ=0
Δ=1
Δ=2
Δ=3
Δ=4
Δ=5
Δ=6
Δ=7
M1 (scaled Manhattan)
1
28.9
31.4
32.3
33.2
34.0
34.6
35.0
36.5
3
21.9
25.6
26.5
27.9
29.2
30.0
30.6
33.1
5
21.0
24.2
25.7
27.0
28.3
29.3
30.7
33.1
M2 (shallow ML)
1
15.8
20.5
22.0
23.3
24.2
24.8
24.7
26.5
3
16.1
20.7
22.1
23.5
24.8
25.2
25.6
26.9
5
17.2
20.6
22.1
23.5
24.5
25.2
25.6
27.0
TABLE X: Full EER (%) by mechanism, enrolment size k , and temporal gap Δ , clipped condition.
Mechanism
Estimator
Δ=0
Δ=1
Δ=2
Δ=3
Δ=4
Δ=5
Δ=6
Δ=7
M1 (scaled Manhattan)
Pooled
21.0
24.2
25.7
27.0
28.3
29.3
30.7
33.1
Particip.-balanced
20.9
24.2
25.6
26.7
28.2
29.4
30.8
32.2
M2 (shallow ML)
Pooled
17.2
20.6
22.1
23.5
24.5
25.2
25.6
27.0
Particip.-balanced
16.2
19.2
20.7
21.9
22.9
23.8
24.2
26.0
M3 (TypeNet-style LSTM)
Pooled
27.2
29.0
30.5
32.1
33.4
34.0
35.2
37.1
Particip.-balanced
27.0
28.6
30.2
31.6
32.8
33.7
34.6
37.7
TABLE XI: EER (%) at k=5 , clipped: pooled-trial vs. per-participant-balanced estimator, all eight values of Δ .
Fold
Mechanism
Seed 0
Seed 1
Seed 2
0
M3
29.7
35.2
29.6
M4
20.6
20.8
21.5
1
M3
31.0
34.5
33.7
M4
18.9
17.7
17.2
2
M3
35.5
32.5
32.6
M4
18.6
19.4
21.0
TABLE XII: Pooled EER (%) for three independently-trained (unfused) seeds of M3 and M4, by test fold.