Few-Shot Learning for Personalised Automated Pain Assessment
Authors: Heinke Hihn, Ibrahim Eisawy, Patrick Thiam, Hans A. Kestler, Friedhelm Schwenker
Organizations: IU International University of Applied Sciences, Berlin, Germany · Institute of Neural Information Processing, Ulm University, Ulm, Germany · Institute of Medical Systems Biology, Ulm University, Ulm, Germany
Pain perception varies substantially across individuals, making it difficult for population-based classifiers to generalise across all subjects in a dataset. One way to account for subject variability is to train personalised classifiers. In this work, we evaluate Few-Shot Learning, a sub-area of Meta-Learning, as an approach to personalisation in automated pain assessment. We re-interpret the shift from population-level to subject-level evaluation as a task-domain shift, where the observed classes remain fixed but the target subject changes. We evaluate our method on the BioVid Pain Database, the SenseEmotion Database, and the PainMonit Experimental Dataset (PMED), reaching 85.75% and 35.49% accuracy on BioVid and 82.37% and 41.88% on SenseEmotion in the binary and multi-class settings under a Leave-One-Subject-Out CV protocol respectively, and 90.47% on PMED, for which only a binary benchmark exists. Using samples to implement k-shot conditioning, the accuracies can be improved to 86.25%, 40.06%, 83.43%, 44.08%, and 91.25%, respectively. To further evaluate the effects and robustness of our method, we provide additional ablation experiments and investigate the personalisation effects. Our results suggest that support-conditioned few-shot adaptation can improve average performance under inter-subject variability.
Figures & tables
Fig 1 : The data in the experiments has been generated by subjecting subjects to various levels of pain (annotated as Ti ) and a baseline of no pain and recording their bodily response by various bio-physical signals such as EDA, EMG, and ECG. This creates time series of signals annotated by the experienced pain level, thus enabling us to train models to predict the pain based on the signals.
Fig 2 : The left figure shows the signal segmentation in the SenseEmotion Database [ 56 ] , where experiments have been carried out on 6.5s windows with a 4s shift from the onset of the elicitation in four pain levels T0,...,T3 . On the right, we show how in the BioVid Database [ 57 ] the windows are of 4.5s with an 4s shift on five pain levels T0,...,T4 . Thus, on sampling rate of 256 Hz we have an input dimensionality 1,664 of and 1,152 , respectively.
Fig 3 : Our approach for Subject-As-A-Task Few-Shot Pain Assessment. First, we sample a task, which is then passed into an encoder for each modality. These streams are combined using a modified Cross Modality Transformer as introduced by Farmani et al. [ 12 ] , as described in Section 3.3 . The resulting support feature maps are used to compute a prototype per class against which we classify the query feature maps using the Cross Alignment Network described in Section 3.4 .
Fig 4 : Our method supports two modes of inference: in the Zero-Shot setting (left), we use a global precomputed bank to initialize the prototypes. Thus, this mode does not require any labelled samples from the subject. In the k -Shot case (right), we use a portion of the held-out subjects samples to initialize the prototypes (the support set) and classify the remainder (the query set). Thus, the Zero-Shot setup is the aggregated population level classifier and the k -Shot setup is the personalised individual classifier.
Method
Binary
Four Classes
Thiam et al. [ 51 ]
83.39%
43.89%
Thiam et al. [ 49 ]
81.05%
40.80%
Thiam et al. [ 54 ]
81.50%
n/a
Thiam et al. [ 53 ]
64.35%
n/a
Kessler et al. [ 32 ]
71.85%
n/a
Ours (Zero-Shot)
82.37% ( ±10.9 )
41.88% ( ±6.6 )
Table 1 : Overview of reported pain recognition performance by algorithm/model on the SenseEmotion dataset. All values are reported for Leave-One-Subject-Out (LOSO) cross-validation. Accuracy is reported with standard deviation in parentheses where available. Bold indicates best Zero-Shot method and underlined indicates that the personalised k -Shot classifiers beats previous methods.
Method
Binary
Five Classes
Farmani et al. [ 12 ]
87.52% ( ±11.0 )
n/a
Aslam et al. [ 4 ]
86.90%
n/a
Li et al. [ 35 ]
86.21%
38.03%
Lu et al. [ 39 ]
85.56%
34.46%
Thiam et al. [ 54 ]
85.32% ( ±13.7 )
n/a
Jiang et al. [ 30 ]
84.58% ( ±13.3 )
39.24% ( ±8.65 )
Table 2 : Overview of reported pain recognition performance by algorithm/model on the BioVid dataset. All values are reported for Leave-One-Subject-Out (LOSO) cross-validation. Accuracy is reported with standard deviation in parentheses where available. Bold indicates best Zero-Shot method and underlined indicates that the personalised k -Shot classifiers beats previous methods.
Method
Binary
Luebke et al. [ 40 ]
93.62%
Gouverneur et al. [ 17 ]
93.26%
Gouverneur et al. [ 17 ]
89.79% ( ± 11.68)
Gouverneur et al. [ 16 ]
87.41% ( ±11.99 )
Gouverneur et al. [ 17 ]
91.09% ( ±9.70 )
Ours (Zero-Shot)
90.47% ( ± 9.23)
Table 3 : Overview of reported pain recognition performance by algorithm/model on the PainMonit dataset. All values are reported for Leave-One-Subject-Out (LOSO) cross-validation. Accuracy is reported with standard deviation in parentheses where available. Bold indicates best Zero-Shot method and underlined indicates that the personalised k -Shot classifiers beats previous methods.
BioVid
Method
T 0 vs. T 1
T 0 vs. T 2
Thiam et al. [ 48 ]
61.15% ( ±12.20 )
66.81% ( ±15.92 )
Gkikas et al. [ 13 ]
52.38%
52.78%
Werner et al. [ 63 ]
48.70%
51.60%
Gkikas et al. [ 13 ]
62.82%
63.68%
Lopez & Picard [ 38 ]
56.44%
59.40%
Table 4 : Comparison of additional binary pain recognition results. Bold indicates best Zero-Shot method and underlined indicates that the personalised k -Shot classifiers beats previous methods.
Fig 5 : Personalisation benefits per subject. A vast majority of 91% to 98% of subjects benefit from personalisation in the BioVid full classification setting, i.e., using part of their samples to initialize the prototypes (see Panels A, B, D). Panel A shows a mean improvement in macro F1 is +0.097 . Panel E shows the distribution of the improvements in the different k settings.
Fig 6 : Personalisation benefits per class on the BioVid dataset. Panels A and B show that the intermediate pain levels 1, 2, and 3 have the largest mean F1 gains from personalisation. Their respective F1 scores improve on average, indicating that personalisation contributes most for pain levels that vary only little in magnitude, compared with the extremes 0 and 4. Panels C and D show how subjects benefit on a per-class basis, indicating the same pattern of improvement in classes 1 to 3.
Fig 7 : Different settings of training for the BioVid (A+B) and the SenseEmotion (C+D) dataset. In all settings increasing the training task size ( k,q values) improves the results with a ceiling on k,q=10 where we do not observe significantly higher results. All significant p -values are <0.001 .
Fig 8 : Mean held-out performance across training and evaluation conditions for binary classification (top) and multi-class classification (bottom) on the SenseEmotion dataset. Columns show accuracy and macro-F1. The Zero-Shot condition (ke,qe)=(0,20) uses learned prototype memory, whereas the remaining conditions use target-subject support.
Fig 9 : Effect of training and evaluation task size on BioVid binary (A+B) and five-class (C+D) performance. Points represent held-out subjects, and boxes show the score distributions. Here, T and E denote the training and configured k -shot evaluation tasks, respectively.
BioVid
SenseEmotion
Modalities
Zero-Shot
k -Shot
Zero-Shot
k -Shot
EMG+RSP
N/A
N/A
63.79% (±11.6)
64.71% (±11.7)
ECG+RSP
N/A
N/A
64.60% (±11.3)
67.51% (±12.2)
EDA+RSP
N/A
N/A
81.64% (±10.4)
82.44% (±10.7)
EDA+EMG
85.52% (±9.31)
85.68% (±9.31)
81.52% (±9.31)
81.99% (±10.7)
ECG+EMG
59.39% (±15.2)
65.14% (±16.0)
58.46% (±10.1)
63.34% (±11.8)
Table 5: Ablation study on the modalities used. In accordance with previous studies [ 39 , 35 ] , EDA+ECG achieves the highest classification accuracy in the binary ( T0 vs. T4 and T0 vs. T3 , respectively) LOSO setting. N/A indicates that the corresponding modality is not available in the dataset. We report accuracy and trained the models with k,q=10 .
Stage
Core configuration
Output
Input
EDA and ECG windows with 1152 steps
B×1152×2
Modality encoders
Independent temporal convolutional frontends with pooling
2×(B×78×16)
CrossMod fusion
Two 8-head self-attention layers followed by bidirectional cross-modal attention
B×78×32
Prototype memory
Mslot=2 learned temporal prototype slots per class
Bt×∣C∣Mslot×78×32
Query encoding
The same modality encoders and CrossMod fusion are applied to each query
Bt×Q×78×32
CAN alignment
Pair-conditioned bidirectional temporal attention and attended cosine similarity, τ=0.025
Bt×Q×∣C∣Mslot
Table 6 : Compact architecture of the proposed CAN model.
Hyperparameter
Setting
Temporal frontend
Filters 8/16 ; kernels 64/16
Pooling / dropout / L2
4,8 / 0.25 / 10−4
CrossMod configuration
2 layers; 8 heads; hidden dim. 128
CAN configuration
τ=0.025 ; meta hidden dim. 32
CAN support source
Learned prototype memory
Prototype memory
2 slots/class; 256 init. samples/class
Table 7 : Main hyperparameters used for the full LOSO experiment.
Department of Computer Science and Engineering, Islamic University of Technology, Gazipur, Bangladesh · Management Information Systems, Metropolitan State University, Minnesota, USA · School of Kinesiology, University of Louisiana at Lafayette, LA, USA