Language as the Interface: Foundation-Model Contrastive Learning Links Transcriptomes and Electrophysiology
Authors: Junbo Shen, Jinying Gao, Bo Lei
Organizations: Department of Computer Science and Engineering, The Chinese University of Hong Kong · State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences, Beijing, China · Beijing Academy of Artificial Intelligence, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides paired measurements of gene expression and intrinsic electrophysiology from the same neuron, establishing a basis for training cross-modal models. Here we introduce LangPatch, a foundation-model-based contrastive learning framework that uses paired Patch-seq data to align pretrained GenePT representations with electrophysiological phenotypes through a language-based interface. Gene descriptions and verbalized electrophysiological profiles are embedded by the same frozen text encoder. A context adapter and projection modules connect the modalities through paired contrastive learning. Across mouse visual, mouse motor, and human cortical cohorts, LangPatch achieves the highest mean transcriptome-to-electrophysiology prediction correlation among the evaluated foundation-model and representation-learning methods. It also improves held-out cross-modal alignment in the two mouse cohorts (FOSCTTM 0.107/0.135 vs. 0.208/0.222 for JAMIE, an existing cross-modal Patch-seq imputation method). It predicts transcriptomic family, type, cortical layer, and marker-gene expression from electrophysiology, exceeding other baselines on most endpoints. More importantly, the method transfers across brain areas and species: a model trained on mouse visual cortex predicts electrophysiology in motor cortex with approximately 70% correlation retention and in human cortex with 47% (58% on acute-slice recordings). Together, these results demonstrate alignment between molecular and functional representations of neurons, providing a building block for multimodal foundation models in neuroscience.
Figures & tables
Figure 1: Language as the interface. (a) Genes enter as text through a frozen embedder; a 0.4M-parameter adapter conditioned on training-cohort statistics adjusts the frozen gene prior, and a cell is the expression-weighted mean of its adapted genes followed by a linear map. (b) Electrophysiology is verbalized and embedded by the same frozen embedder, then linearly mapped. (c) A symmetric InfoNCE objective on paired cells aligns the two; the embedder is frozen. (d) The shared space has numerous downstream applications: predicting electrophysiological features from the transcriptome (Sec. 5.1 ), reading transcriptomic family, type, layer and marker genes from electrophysiology alone (Sec. 5.2 ), zero-shot transfer across cortical areas and species (Sec. 5.3 ), and gene-level attribution of each feature (Sec. 6 ).
Mouse visual (39 feat.)
Mouse motor (29 feat.)
Human Lee (75 feat.)
Method
r↑
AUROC ↑
MAE ↓
r↑
AUROC ↑
MAE ↓
r↑
AUROC ↑
MAE ↓
Ours
0.522 ± 0.018
0.779 ± 0.008
8.98 ± 0.12
0.520 ± 0.021
0.801 ± 0.005
7.72 ± 0.33
0.414 ± 0.030
0.725 ± 0.014
9.82 ± 0.27
Plain GenePT MLP
0.505 ± 0.016
0.768 ± 0.005
9.22 ± 0.04
0.502 ± 0.020
0.790 ± 0.002
8.02 ± 0.37
0.387 ± 0.021
0.713 ± 0.009
9.90 ± 0.19
Geneformer MLP
0.492 ± 0.017
0.760 ± 0.005
9.54 ± 0.16
0.146 ± 0.028
0.574 ± 0.010
10.44 ± 0.32
0.401 ± 0.034
0.719 ± 0.010
9.88 ± 0.27
UCE MLP
0.410 ± 0.016
0.714 ± 0.007
10.51 ± 0.14
0.462 ± 0.014
0.762 ± 0.009
8.57 ± 0.38
0.312 ± 0.035
0.670 ± 0.018
10.65 ± 0.37
Nicheformer MLP
0.486 ± 0.016
0.756 ± 0.005
9.56 ± 0.08
0.503 ± 0.018
0.784 ± 0.007
8.06 ± 0.27
0.386 ± 0.028
0.710 ± 0.008
10.13 ± 0.24
Table 1: Forward prediction on held-out cells (mean ± s.d. over five folds): per-feature Pearson r , AUROC and raw MAE. Ours: context-adapted GenePT with alignment and per-feature heads whose width is selected on the validation fold (Appendix A ).
Figure 2: Shared space. (a) Reverse readout of transcriptomic family from electrophysiology (macro-F1). (b) Held-out FOSCTTM (lower is better; 0.5 is chance) and (c) family label-transfer accuracy for the methods that expose a joint space, JAMIE re-run on our folds. (d) Training time and peak GPU memory comparisons for one fold, cost-efficiency (Appendix C )
Figure 3: Zero-shot transfer without target training. (a) Per-concept Pearson correlations for visual ↔ motor transfer, evaluated using the cross-area matched electrophysiology features and permutation-based chance levels. (b) Correlation retention, defined relative to the corresponding in-domain performance, for cross-area and mouse-to-human transfer. (c,d) Few-shot calibration as a function of the number of paired target cells for cross-area and mouse-to-human transfer.
Figure 4: Interpretability of the shared space. (a, b) UMAP of the aligned space for the held-out cells of one fold: transcriptome embeddings (circles) and electrophysiology embeddings (triangles), coloured by transcriptomic family; grey segments join a random subset of pairs. (c, d) Gene × feature attribution of the forward heads (gradient × input, mean over five folds, z -scored per feature) for the fifteen most attributed genes, feature columns grouped by physiological family.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Cohort
Cells
Genes
Ephys features
Labels
Role in the paper
Mouse visual cortex (Gouwens 2020)
3,654
1,293
39
6 families, 4 layers
In-domain benchmark; source cohort of every zero-shot transfer
Mouse motor cortex (Scala 2021)
1,208
1,277
29
9 families, 4 layers
In-domain benchmark; cross-area transfer, both directions
Table 2: Cohorts. In-domain benchmarks use five-fold cross-validation; zero-shot targets are scored through pre-registered feature crosswalks against measured electrophysiology.
Component
Setting
Stage 1: context adapter and alignment
Context vector C
24 (visual), 30 (motor), 18 (human): 4 global + per-family and per-layer mean and detection rate
adapter 0.40M; Pcell and Pephys 9.45M each; auxiliary regression head 9.56M
Contrastive loss
symmetric InfoNCE, temperature τ=0.07
Appendix
Table 3: Hyperparameters and trainable parameter counts of the two stages.
Method
FOSCTTM ↓
LTA family ↑
LTA type ↑
LTA layer ↑
Silhouette ↑
Mouse visual
Ours
0.107 ± 0.004
0.926 ± 0.032
0.298 ± 0.020
0.462 ± 0.012
0.178 ± 0.040
JAMIE
0.208 ± 0.006
0.777 ± 0.033
0.207 ± 0.022
0.407 ± 0.026
0.102 ± 0.020
GenePT–ephys VAE
0.265 ± 0.136
0.629 ± 0.189
0.113 ± 0.064
0.301 ± 0.138
0.124 ± 0.175
Nicheformer + CLIP
0.187 ± 0.004
0.850 ± 0.017
0.239 ± 0.028
0.434 ± 0.025
0.057 ± 0.007
Mouse motor
Appendix
Table 4: Integration metrics on held-out cells (FOSCTTM, label-transfer accuracy, silhouette) and aligner cost.
Mouse visual
Mouse motor
Method
Family F1
Type F1
Layer F1
Marker r
Family F1
Type F1
Layer F1
Marker r
Ours (context-guided multitask head)
0.779 ± 0.028
0.382 ± 0.017
0.459 ± 0.019
0.539 ± 0.007
0.710 ± 0.075
0.572 ± 0.043
0.575 ± 0.035
0.572 ± 0.009
Direct supervised ephys
0.755 ± 0.018
0.397 ± 0.019
0.413 ± 0.017
0.453 ± 0.031
0.678 ± 0.047
0.557 ± 0.083
0.486 ± 0.020
0.509 ± 0.014
Ephys → GenePT PCA-ridge
0.665 ± 0.028
0.183 ± 0.031
0.404 ± 0.016
0.456 ± 0.032
0.625 ± 0.034
0.380 ± 0.055
0.477 ± 0.040
0.509 ± 0.015
Ephys → Geneformer MLP
0.745 ± 0.023
0.320 ± 0.026
0.429 ± 0.034
0.450 ± 0.111
0.363 ± 0.220
0.210 ± 0.154
0.313 ± 0.102
0.216 ± 0.154
Ephys → UCE MLP
0.659 ± 0.042
0.150 ± 0.007
0.344 ± 0.043
0.391 ± 0.021
0.672 ± 0.065
0.383 ± 0.042
0.496 ± 0.036
0.463 ± 0.030
Appendix
Table 5: Reverse readouts on held-out cells (mean ± s.d. over five folds): macro-F1 for family, fine type and layer; Pearson r for marker genes and programs.
Concept
Visual feature (IPFX)
Motor feature (Scala)
Sign certain
resting potential
vrest
Resting.membrane.potential..mV.
yes
input resistance
ri
Input.resistance..MOhm.
yes
membrane time constant
tau
Membrane.time.constant..ms.
yes
sag
sag
Sag.ratio
no
rheobase
threshold_i_long_square
Rheobase..pA.
yes
AP threshold
threshold_v_long_square
AP.threshold..mV.
yes
Appendix
Table 6: Pre-registered visual ↔ motor crosswalk.
Method
Pearson r
AUROC
scaled MSE
raw MAE
Δ vs. GenePT
feat. better
p (feat.)
p (fold)
Ours (head selected on validation)
0.414 ± 0.030
0.725
1.182
9.82
+0.027
48/75
7.2×10−5
0.062
Ours (mouse head config., no selection)
0.400 ± 0.024
0.722
1.221
10.00
+0.013
44/75
0.051
0.31
Mouse stage-1 frozen + human heads
0.395 ± 0.032
0.722
1.225
10.03
+0.008
36/75
0.55
0.62
Plain GenePT MLP
0.387 ± 0.021
0.713
1.196
9.90
ref.
—
—
—
Geneformer MLP
0.401 ± 0.034
0.719
1.206
9.88
+0.014
42/75
0.13
0.31
UCE MLP
0.312 ± 0.035
0.670
1.265
10.65
-0.075
11/75
1.7×10−10
0.062
Appendix
Table 7: Human-Lee in-domain benchmark (612 cells, 75 features) and the nested-validation head selection.
Representation (+ ridge head)
n=25
n=50
n=100
n=200
all
Mouse stage 1, frozen
0.270
0.297
0.329
0.367
0.412
Plain GenePT
0.197
0.264
0.333
0.374
0.420
Geneformer
0.195
0.256
0.305
0.347
0.374
Human stage 1 (saw all pairs; reference only)
0.306
0.338
0.369
0.400
0.431
Appendix
Table 8: Low-resource human forward prediction: mean Pearson r over 75 features with a ridge head trained on n labelled Lee 2023 cells (median over 5 folds × 10 draws; “all” uses the full training fold, 5 folds). The last row’s stage 1 saw the electrophysiology of all training cells and is a reference, not a low-resource result.
Mouse visual cortex
Criterion
Ours
Plain GenePT MLP
JAMIE
null
Train-fold markers: recovery AUC ( K≤400 )
0.57 ± 0.01
0.57 ± 0.02
0.06 ± 0.03
0.16 μ
Train-fold markers among the top 200 (fraction)
0.64 ± 0.02
0.64 ± 0.03
0.05 ± 0.03
0.21
Curated inhibitory markers among the top 200 (fraction of 11)
0.40 ± 0.05 (4.4/11)
0.60 ± 0.05 (6.6/11)
0.00 ± 0.00 (0.0/11)
0.36
Train-fold markers: fold enrichment, top 50
6.60 ± 0.34
6.60 ± 0.44
0.35 ± 0.33
1.78
Inhibitory markers: fold enrichment, top 50
9.40 ± 0.00
8.93 ± 1.97
0.00 ± 0.00
4.70
Appendix
Table 9: Attribution benchmark: the proposed model, the Plain GenePT MLP and JAMIE, attributed with the same gradient × input protocol on the same held-out cells and folds, scored on biological plausibility, stability across folds, faithfulness under gene deletion and keep-only, and gene removal in the joint space.
Target
cells / concepts
Method
mean r
> null
R (verdict)
p
Visual → motor
1,208 / 11
Ours
0.344
11/11
0.70 (A)
0.17
Plain GenePT
0.317
11/11
0.66 (A)
ref.
Motor → visual
3,654 / 11
Ours
0.384
11/11
0.71 (A)
0.46
Plain GenePT
0.371
11/11
0.69 (A)
ref.
Lee 2023, all
704 / 25
Ours
0.282
18/25
0.47 (B)
0.10
Plain GenePT
0.271
18/25
0.44 (B)
ref.
Appendix
Table 10: Zero-shot transfer summary across cortical areas and species: mean r , concepts above the permutation null with the registered sign, retention and verdict.
Department of Brain and Cognitive Engineering, Korea University, Seoul, Republic of Korea · Department of Artificial Intelligence, Korea University, Seoul, Republic of Korea