scTrilemma: Balancing Identity, Invariance, and Fidelity in Single-Cell Representation Learning
Organizations: KAIST · Work done during an internship at HITS · HITS
Abstract
Single-cell RNA-seq representation learning is fundamentally label-free: cell identities, states, and contexts are not fixed training targets, so what constitutes signal or nuisance is analysis-dependent. A single representation must therefore preserve biological identity and state, remain robust to nuisance context, and retain the gene-level variation needed for expression analysis, three demands we call the representation trilemma. To tackle this problem, we introduce scTrilemma, a latent-bottleneck VAE that routes expression-derived variation to the embedding, the decoder, or the prior rather than forcing all of it through one embedding. It gates gene tokens by expression, routes the cell representation through the decoder, and conditions the prior on unlabeled pseudo-bulk context, under a single reconstruction objective and without target annotations or auxiliary representation losses. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, scTrilemma leads all three demands at once and preserves biological-state, differential-expression, and pathway structure across multiple disease settings. Latent interventions further show that context can be removed at almost no cost to the other demands, leaving identity against fidelity as the remaining tension. Code is publicly available at https://github.com/yunhak0/scTrilemma.
Figures & tables
| Biological identity | Context invariance | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | NMI | ARI | ASW | cLISI | Iso. Label | BRAS | iLISI | PCR |
| scVI | ||||||||
| Geneformer | ||||||||
| scGPT | ||||||||
| CellPLM | ||||||||
| scPRINT | ||||||||
| Metric | scVI | Geneformer † | scGPT † | CellPLM | scPRINT ‡ | scTrilemma |
|---|---|---|---|---|---|---|
| Biological identity and state | ||||||
| Cell-type BA chance | ||||||
| Neighbor purity | ||||||
| Disease-state BA | ||||||
| Context invariance | ||||||
| Residual donor invariance | ||||||
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Cohort | Dataset ID prefix | Fixed context | Disease/state (vs. normal) | Evaluation support † |
|---|---|---|---|---|
| Kidney | 867757c1 | 10x multiome, kidney | obstructive nephropathy | 13 / 2,243 / 520 / 1,524 |
| Liver | e3ed2ba4 | BD Rhapsody WTA, liver | metastatic colorectal carcinoma | 11 / 2,073 / 440 / 2,135 |
| Brain | 203025fe | 10x multiome, dorsolateral prefrontal cortex | cognitive disorder | 7 / 830 / 280 / 1,821 |
| Tendon | acd544d0 | 10x 3’ v3, quadriceps femoris tendon | injury | 4 / 854 / 160 / 1,555 |
| Colon | 829a3cd1 | 10x 3’ v3, sigmoid colon | colorectal cancer | 6 / 959 / 240 / 1,639 |
| Dataset ID prefix | Disease (vs. normal) | Fixed context | Evaluation support † |
|---|---|---|---|
| 16023185 | Colon adenocarcinoma | 10x 3 ′ v2, colorectum | 43 / 38 / 12,000 |
| f14bc322 | Interstitial lung disease | 10x 5 ′ v2, lung | 93 / 33 / 12,000 |
| d68a8b48 | Bronchopulmonary dysplasia | 10x 3 ′ v3, middle lobe of right lung | 21 / 35 / 12,000 |
| 19053a82 | Ulcerative colitis | 10x 3 ′ v2, colon | 11 / 40 / 12,000 |
| Candidate collection | Shared donors | Eligible strata | Evaluation support † | Outcome |
|---|---|---|---|---|
| RPE/choroid atlas | 4 | 24 | 4 / 9 | Selected |
| Anterior-segment atlas | 8 | 3 | 2 / 3 | Lower matched support |
| Ocular-surface atlas | 0 | 0 | 0 / 0 | No shared donor |
| Donor | Eligible cell types | Strata | Cells per repeat |
|---|---|---|---|
| BCM_22_0500 | endothelial cell of venule, fenestrated endothelial cell, fibroblast, macrophage, melanocyte | 5 | 878 |
| BCM_22_0698 | endothelial cell of venule, fenestrated endothelial cell, fibroblast, macrophage, melanocyte, pericyte | 6 | 1,068 |
| BCM_22_0769 | T cell, endothelial cell of venule, fibroblast, macrophage, melanocyte, retinal pigment epithelial cell | 6 | 1,120 |
| BCM_22_0784 | T cell, endothelial cell of venule, fibroblast, macrophage, melanocyte, monocyte, retinal pigment epithelial cell | 7 | 1,290 |
| Total | 9 unique cell types | 24 | 4,356 |
| Dataset ID prefix | Disease (vs. normal) | Evaluation support † |
|---|---|---|
| 30c2a6fd | Open-angle glaucoma | 15 / 6 / 1,230 |
| 84cfa5aa | Intestinal failure-associated liver disease | 14 / 6 / 2,045 |
| 867757c1 | Obstructive nephropathy | 13 / 6 / 1,524 |
| e3ed2ba4 | Metastatic colorectal carcinoma | 11 / 6 / 2,327 |
| 203025fe | Alzheimer disease | 8 / 6 / 1,877 |
| 829a3cd1 | Colorectal cancer | 8 / 6 / 1,811 |
| Package | Version |
|---|---|
| torch | 2.8.0+cu128 |
| lightning / pytorch-lightning | 2.6.0 |
| hydra-core / omegaconf | 1.3.2 / 2.3.0 |
| numpy / scipy / pandas | 2.3.5 / 1.16.3 / 2.3.3 |
| anndata / scanpy | 0.12.7 / 1.11.5 |
| scikit-learn | 1.8.0 |
| Model | Params | Emb. dim | Batch size | Median time (s) | Cells/s | Peak GPU memory |
|---|---|---|---|---|---|---|
| scVI | 8.0M | 50 | 512 | 0.041 | 49,225.2 | 0.09GB |
| scGPT | 51.9M | 512 | 64 | 2.141 | 934.0 | 1.52GB |
| CellPLM | 65.8M | 512 | 2,000 † | 0.396 | 5,056.6 | 0.88GB |
| Geneformer | 316.3M | 1152 | 16 | 32.349 | 61.8 | 3.15GB |
| scTrilemma | 58.3M | 512 | 256 | 6.433 | 310.9 | 19.79GB |
| Category | Setting | Value |
|---|---|---|
| Training corpus | CELLxGENE Census release | 20250130 |
| Zero-shot evaluation release | CELLxGENE Census release | 20251108 |
| Held-out zero-shot benchmark | Release-based benchmark | 89 newly appearing datasets |
| Counterfactual subset | Held-out blood/protocol subset | 10 blood datasets from the 20251108 release |
| Random seed | Seed | 42 |
| Batch size | Cells per batch | 256 per GPU (1,024 total across 4 GPUs) |
| Category | Setting | Value |
|---|---|---|
| Optimizer | AdamW | learning rate , weight decay , betas |
| Scheduler | Cosine with linear warmup | 1,250 warmup steps (5% of 25,000), then cosine decay to step 25,000 |
| KL loss | KL coefficient, warmup | , 2,500 steps (10% of the 25,000-step recipe) |
| Trainer | Precision, devices, strategy | bf16-mixed; 4 GPUs; DDP |
| Trainer schedule | Steps, validation interval | 25,000, 5,000 |
| Gradient clipping | Clip value | 1.0 |
| Component | Setting | Role |
|---|---|---|
| Model family | Latent-bottleneck VAE with cross-attention encoder and decoder | Compresses gene-expression tokens into latent cell tokens before reconstruction. |
| Trainable parameters | 58,347,460 | Counted from the active final model configuration. |
| Gene vocabulary | 61,890 Ensembl genes | Gene IDs are indexed by the 20250130 CELLxGENE vocabulary. |
| Token width | Shared width for gene tokens, latent tokens, attention blocks, and decoder states. | |
| Gene embedding | Learnable, 512-dimensional | Provides the base gene identity token. |
| Expression-gated gene encoding | Multiplicative expression gate | Gates gene embeddings by measured expression before compression. |
| Component | Setting | Role |
|---|---|---|
| Grouping unit | Dataset-donor group | Defines the collection-level unit for pseudo-bulk profiles. |
| Pseudo-bulk profile | Per-cell CP10K normalization, group mean, then elementwise | Produces the group-level expression signature used for conditioning. |
| Centroid construction | k-means with , random seed 42, and five random initializations | Learns expression-derived collection archetypes from pseudo-bulk profiles. |
| Soft assignment | Euclidean distance to centroids with distance standardization | Produces a 32-dimensional soft pseudo-bulk code. |
| Zero-shot-release assignment | Fixed 20250130 centroids; no centroid refit on 20251108 | Assigns unseen release groups by soft assignment to the precomputed centroids. |
| Prior conditioning | Diagonal Gaussian prior | Conditions the VAE prior using the pseudo-bulk code. |
| Evaluation | Metrics | Reported on | Direction † |
|---|---|---|---|
| RQ1 biological identity | NMI, ARI, ASW-label, cLISI, isolated labels | Cell embedding | |
| RQ1 context invariance | BRAS, iLISI, PCR comparison | Cell embedding | |
| RQ2 biological identity and state | Cell-type BA minus chance, neighbor purity, disease-state BA | Cell embedding | |
| RQ3 context invariance | Residual donor invariance, conditional donor BRAS, local donor mixing | Cell embedding | |
| RQ4 expression fidelity | logFC Spearman, DEG Jaccard and sign, pathway Jaccard and score Spearman | Reconstructed expression | |
| Supplementary identity and state | Program-neighbor Spearman, purity, accuracy, balanced accuracy | Embedding neighborhoods and expression |
| Marker program | Share | scVI | Geneformer | scGPT | CellPLM | scPRINT | scTrilemma |
|---|---|---|---|---|---|---|---|
| CD4 T cell | 19.3 | 0.8832 | 0.8983 | 0.8869 | 0.8820 | 0.8575 | 0.9070 |
| CD8/cytotoxic T | 15.0 | 0.8411 | 0.8581 | 0.8256 | 0.8032 | 0.7853 | 0.8468 |
| B cell | 13.5 | 0.5883 | 0.5850 | 0.5730 | 0.5752 | 0.5511 | 0.5885 |
| NK cell | 11.7 | 0.7811 | 0.7819 | 0.7713 | 0.7637 | 0.7686 | 0.7811 |
| Classical monocyte | 5.8 | 0.6446 | 0.6518 | 0.6554 | 0.6457 | 0.6679 | 0.6656 |
| Non-classical mono. | 3.8 | 0.7471 | 0.7620 | 0.7548 | 0.7422 | 0.7271 | 0.7605 |
| Cohort | Metric | scVI | Geneformer | scGPT | CellPLM | scPRINT | scTrilemma |
|---|---|---|---|---|---|---|---|
| Colon | Purity | 0.6432 | 0.6371 | 0.6312 | 0.6310 | 0.6249 | 0.6583 |
| Accuracy | 0.7276 | 0.7267 | 0.7274 | 0.7093 | 0.7299 | 0.7507 | |
| Balanced acc. | 0.6499 | 0.6511 | 0.6496 | 0.6220 | 0.6482 | 0.6733 | |
| ILD | Purity | 0.6083 | 0.6082 | 0.5994 | 0.5952 | 0.5962 | 0.6385 |
| Accuracy | 0.6920 | 0.6941 | 0.6868 | 0.6831 | 0.6830 | 0.7359 | |
| Balanced acc. | 0.5802 | 0.5828 | 0.5750 | 0.5659 | 0.5667 | 0.6294 |
| Metric | scVI | Geneformer | scGPT | CellPLM | scPRINT | scTrilemma |
|---|---|---|---|---|---|---|
| Modality | ||||||
| Modality ASW | ||||||
| Conditional iLISI |
| Cohort | Method | logFC Spearman | top-100 DEG Jaccard | top-100 DEG Sign | Pathway Jaccard | Pathway score Spearman |
|---|---|---|---|---|---|---|
| OAG | scVI | 0.1573 | 0.1379 | 0.6367 | 0.2406 | -0.1662 |
| scPRINT | 0.5559 | 0.2163 | 0.7550 | 0.2535 | -0.1124 | |
| CellPLM | 0.2194 | 0.1655 | 0.7000 | 0.2594 | -0.1456 | |
| scTrilemma | 0.4302 | 0.2498 | 0.8683 | 0.3261 | 0.1589 | |
| IFALD | scVI | 0.1932 | 0.1401 | 0.7233 | 0.2247 | -0.1954 |
| scPRINT | 0.4062 | 0.2528 | 0.7600 | 0.3823 | 0.2449 |
| Method | Pearson | Spearman | MAE | MSE |
|---|---|---|---|---|
| All shared genes | ||||
| scVI | 0.3722 0.0846 | 0.3098 0.0602 | 1.0041 0.1157 | 1.5788 0.3013 |
| scPRINT | 0.4751 0.0577 | 0.3962 0.0617 | 1.0475 0.2220 | 1.6569 0.5433 |
| CellPLM | 0.5785 0.1045 | 0.4404 0.0759 | 0.8135 0.2247 | 1.3003 0.5483 |
| scTrilemma | 0.6380 0.0863 | 0.4524 0.0584 | 1.4162 0.2310 | 2.2828 0.6831 |
| Genes detected in 20% of cells | ||||
| NMI | Label ASW | BRAS | iLISI | Recon. Pearson | MSE | |
| Push identity (centroid shrinkage) | ||||||
| (40/89) | (86/89) | (0/75) | (65/75) | (0/89) | (88/89) | |
| (46/89) | (85/89) | (0/75) | (69/75) | (0/89) | (89/89) | |
| (46/89) | (77/89) | (0/75) | (69/75) | (0/89) | (89/89) | |
| Push identity (label centroids) | ||||||
| (89/89) | (89/89) | (1/75) | (61/75) | (7/89) | (89/89) | |
| NMI | ASW | BRAS | iLISI | HVG Pearson | MSE | |
|---|---|---|---|---|---|---|
| Backbone | ||||||
| E-Gate only | ||||||
| E-Gate + C-Route (full model) |
| NMI | ASW | BRAS | iLISI | Recon. Pearson | MSE | |
|---|---|---|---|---|---|---|
| scVI, no batch covariate | ||||||
| + decoder batch covariate | ||||||
| + PB-conditioned prior |
| Matched | N = 50 | Rare-80% | Dominant-80% | Global shuffle | Zero code | |
|---|---|---|---|---|---|---|
| KL (%) | ||||||
| Mean | ||||||
| Max outputs | – |
| Metric | Full | Shuffled codes | w/o PB-Cond |
|---|---|---|---|
| NMI | |||
| Label ASW | |||
| BRAS | |||
| iLISI | |||
| HVG Pearson | |||
| MSE |
| Module | Top genes | Source-backed enriched terms | Descriptive label |
|---|---|---|---|
| 54 | PTPRD , NPAS3 , NRXN1 , NLGN1 , PCDH9 , LSAMP | Brain and oligodendrocyte-precursor cell-marker terms; immature-neuron terms; neuron projection and adhesion genes | Neuronal projection / adhesion |
| 22 | MT-CO1 , MT-CO2 , MT-ATP6 , MT-ND2 , CALM2 , B2M | rRNA processing; metabolism of RNA; mitochondrial RNA processing and degradation | Translation / mitochondrial RNA axis |
| 16 | EDIL3 , NCAM2 , SLC1A2 , PTN , PLP1 , PTPRZ1 | Oligodendrocyte and oligodendrocyte-progenitor marker terms; astrocyte and radial-glia terms | Oligodendrocyte / glial |
| 1 | HLA-B , PTPRC , CD83 , CCL4 , MT-ATP8 , MT-ND4L | Dendritic-cell, NK-cell, and T-cell marker terms; IL-2/STAT5 and TNF-alpha signaling | Immune / inflammatory |