When to Unpair: Regulating Pairing Dependence in Medical Visual In-Context Learning
Organizations: Cornell University · Weill Cornell Medicine
Abstract
Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, while the labels collectively indicate the requested task. We diagnose dependence on individual pairings with a test-time derangement that reassigns every support label to another support image while preserving the query, support images, and label multiset. The resulting pairing gap, defined as shuffled-minus-matched performance, shows that all four released models depend on the pairing, to widely varying degrees. Further analysis of a paired-trained model reveals support-associated spurious regions and lesion-size biases even with real, unaltered supports, alongside sensitivity to mis-registered support labels. To regulate this dependence, we introduce a late unpairing curriculum (LUC), which starts with matched training and then applies random unpairing, replacing each support label with that of another support in the same episode. LUC nearly closes the pairing gap on two backbones while maintaining or improving matched-support performance across all evaluated task types, with gains extending to held-out tasks and cross-dataset episodes. It also mitigates these failure modes. On BraTS whole-tumor segmentation, matched-support DSC rises from 0.733 to 0.857 while the gap shrinks from -0.184 to -0.008. In a released model, brief fine-tuning with random unpairing reduces the gap. A reversed curriculum that places the same number of unpairing epochs at the start of training leaves a large gap. This shows that pairing dependence is shaped by the order of training and not only by the amount of unpaired training.
Figures & tables
| Whole tumor | Edema | ||||
| Backbone | Fusion | matched | gap | matched | gap |
| 2D | |||||
| UniverSeg ( Butoi et al., 2023 ) | averaging | 0.731 | 0.508 | ||
| SegGPT ( Wang et al., 2023b ) | feature ensemble | 0.713 | 0.510 | ||
| 3D | |||||
| Neuroverse3D ( Hu et al., 2025 ) | mean-over- | 0.846 | 0.593 | ||
| paired ( ) | LUC 0.5 | LUC 0.2 | ||||
|---|---|---|---|---|---|---|
| Task (metric) | matched | gap | matched | gap | matched | gap |
| Segmentation (DSC ) | ||||||
| Translation (PSNR ) | ||||||
| Anatomical seg. (DSC ) | ||||||
| Inpainting (PSNR ) | ||||||
| Bias corr. (PSNR ) | ||||||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Ingredient | Value (fixed across the family) |
|---|---|
| Architecture | Neuroverse3D, channels , 71M params |
| Data mix | BraTS 2021 three T1 datasets (ABCD, ADNI, OASIS-3) |
| Tasks | segmentation, translation, anatomical seg., inpainting, bias correction |
| Optimizer | Adam, lr , no weight decay |
| Learning-rate schedule | held fixed across the family |
| Loss |
| model | mirror | ||
| paired | for all | 0 | none |
| unpaired | for all | 1 | none |
| EUC 0.5 | for , else | 0.5 | LUC 0.5 |
| LUC 0.5 | for , else | 0.5 | EUC 0.5 |
| EUC 0.2 | for , else | 0.2 | LUC 0.2 |
| LUC 0.2 | for , else | 0.2 | EUC 0.2 |
| model | matched | cross-patient | zero | noise | shuffled |
|---|---|---|---|---|---|
| paired | 0.733 | 0.661 | 0.655 | 0.461 | 0.549 |
| unpaired | 0.611 | 0.713 | 0.714 | 0.573 | 0.719 |
| LUC 0.5 | 0.857 | 0.832 | 0.826 | 0.710 | 0.848 |
| LUC 0.2 | 0.878 | 0.878 | 0.879 | 0.808 | 0.875 |
| EUC 0.5 | 0.731 | 0.669 | 0.631 | 0.541 | 0.640 |
| EUC 0.2 | 0.731 | 0.646 | 0.650 | 0.534 | 0.601 |
| model | Seg DSC | Transl PSNR | AnatSeg DSC | Inp PSNR | Bias PSNR |
|---|---|---|---|---|---|
| paired | 0.733 / | 22.406 / | 0.704 / | 17.517 / | 26.588 / |
| unpaired | 0.611 / | 22.146 / | 0.685 / | 17.327 / | 25.958 / |
| EUC 0.5 | 0.731 / | 22.389 / | 0.708 / | 17.339 / | 25.661 / |
| LUC 0.5 | 0.857 / | 22.557 / | 0.806 / | 17.756 / | 27.447 / |
| EUC 0.2 | 0.731 / | 22.486 / | 0.694 / | 17.651 / | 26.516 / |
| LUC 0.2 | 0.878 / | 22.698 / | 0.791 / | 17.878 / | 26.338 / |
| Paired ( ) | LUC 0.5 | LUC 0.2 | EUC 0.5 | EUC 0.2 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Cohort | matched | gap | matched | gap | matched | gap | matched | gap | matched | gap |
| ABCD (children) | ||||||||||
| ADNI (elderly) | ||||||||||
| OASIS-3 (mixed) | ||||||||||
| backbone | training | matched DSC | pairing gap (DSC) |
|---|---|---|---|
| Neuroverse3D ( Hu et al., 2025 ) (3D) | paired | 0.733 | |
| LUC 0.5 | 0.857 | ||
| UniverSeg ( Butoi et al., 2023 ) (2D) | paired | 0.590 | |
| LUC 0.5 | 0.602 |
| model | correct | deranged | rolled16 |
|---|---|---|---|
| paired | 0.754 | 0.585 | 0.685 |
| unpaired | 0.630 | 0.741 | 0.710 |
| EUC 0.5 | 0.755 | 0.667 | 0.686 |
| EUC 0.2 | 0.751 | 0.650 | 0.699 |
| LUC 0.5 | 0.860 | 0.848 | 0.852 |
| LUC 0.2 | 0.880 | 0.879 | 0.880 |
| task | data | input | paired ( ) | LUC 0.5 | LUC 0.2 | ||||
| matched | gap | matched | gap | matched | gap | ||||
| Segmentation (DSC) | |||||||||
| Hippocampus † | T1 cohorts | 60 | – | 0.608 | 0.011 | 0.706 | 0.004 | 0.707 | 0.017 |
| Caudate | T1 cohorts | 60 | – | 0.632 | 0.014 | 0.720 | 0.004 | 0.716 | 0.014 |
| Thalamus † | T1 cohorts | 60 | – | 0.758 | 0.001 | 0.853 | 0.000 | 0.832 | 0.002 |
| Edema from FLAIR † | BraTS | 40 | – | 0.627 | 0.088 | 0.695 | 0.003 | 0.711 | 0.004 |