Fluorescence microscopy reveals where proteins localize, but only a limited number of proteins can be imaged in the same cell; generating these images from amino-acid sequence and the cell's morphological context enables in silico localization of unimaged proteins. The two conditions, however, play asymmetric roles: morphological context is spatially aligned with the target, whereas sequence is non-spatial and must specify protein-dependent localization within it, with recurring coarse patterns shared across proteins and finer protein-specific variation. Existing generators condition on both jointly, without separating what each explains. We introduce FACET (Factorized Asymmetric Conditioning for Efficient Transport), a probabilistic generative framework that encodes this structure as an explicit inductive bias: sequence semantics are learned from what context leaves unexplained, coarse localization regularities are shared across proteins through a semantic memory, and protein-specific variation is a bounded residual around them. A variance-preserving state projection further lets FACET perform continuous stochastic transport through a pretrained diffusion predictor with minimal parameter overhead. On held-out proteins, FACET improves spatial overlap by 34.3% on the Human Protein Atlas and 14.0% on OpenCell over a backbone-matched baseline, and reduces FID by 27.2% and 46.5%, respectively, with 75% fewer network evaluations. It also substantially improves protein-association structure recovery and yields better-calibrated predictions, while detailed ablations show complementary contributions from its design choices. These results identify factorized asymmetric conditioning, rather than generator capacity alone, as a key lever for high-fidelity, efficient, and biologically meaningful cellular image synthesis.
Figures & tables
Figure 1: Probabilistic graphical model for FACET .
Figure 2: Representative held-out HPA examples: morphology context, target, and model outputs. CELL-E2 uses nucleus-only context; CELL-Diff and FACET use all three context channels.
Method
NFE ↓
MI ↑
MS-SSIM ↑
Haralick ↑
PSNR ↑
IoUcont↑
FID ↓
CELL-E2
–
0.4041
0.1370
0.5644
6.36
0.1755
166.4
CELL-Diff
100
0.5355
0.3471
0.5182
15.77
0.2478
45.60
FACET
25
0.5481
0.3651
0.5854
16.15
0.3328
33.20
Table 1: Comparison on the HPA test set. Higher is better for all metrics except NFE and FID. Best results are shown in bold.
Context
System
MI ↑
MS-SSIM ↑
Haralick ↑
PSNR ↑
IoUcont↑
FID ↓
Nucleus
CELL-Diff
0.3091
0.2777
0.4932
15.32
0.2047
51.10
FACET
0.2519
0.1712
0.5010
14.66
0.2768
52.12
Nucleus + ER
CELL-Diff
0.4562
0.3154
0.4898
15.52
0.2182
60.00
FACET
0.5120
0.3413
0.5726
15.96
0.3268
33.83
Nucleus + MT
CELL-Diff
0.5097
0.3404
0.5196
15.60
0.2435
47.60
FACET
0.5226
0.3524
0.5796
16.03
0.3266
33.24
Table 2: Performance across cellular-context compositions on HPA. All systems use the same proteins, sequence inputs, and sampling protocol per composition. Bold marks the better value within each composition.
Method
NFE ↓
MI ↑
MS-SSIM ↑
Haralick ↑
PSNR ↑
IoU cont ↑
FID ↓
CELL-E2
–
0.1649
0.1020
0.4542
11.31
0.1779
245.8
CELL-Diff
100
0.2732
0.4078
0.3663
13.20
0.3722
68.90
FACET (HPA-adapted)
25
0.2753
0.3653
0.5177
13.85
0.4222
43.33
FACET (OpenCell-adapted)
25
0.2998
0.3893
0.5199
13.75
0.4244
36.83
Table 3: Comparison on the OpenCell test set (nucleus-only context). All systems except CELL-E2 share the OpenCell-trained CELL-Diff backbone; HPA-adapted transfers the FACET modules from HPA without refitting, whereas OpenCell-adapted fits them on OpenCell. Best results are in bold.
System
STRING score threshold
Pairs (assoc./null)
Pearson Δ↑
passoc.vs.null
Baseline
0.85
47/2659
+0.1005
0.298
FACET
0.85
47/2659
+0.1875
1.3×10−3
Baseline
0.70
65/2641
+0.0954
0.170
FACET
0.70
65/2641
+0.1718
4.5×10−4
Baseline
0.50
112/2594
+0.0922
0.042
FACET
0.50
112/2594
+0.1369
2.5×10−4
Table 4: Spatially resolved protein-association pattern recovery under a localization-matched evaluation adapted from ( Sun et al., 2026 ) . The baseline denotes CELL-Diff. Pearson similarity is computed after removing the per-cell mean image. Δ measures the separation between STRING-associated and matched non-associated pairs; positive values indicate stronger recovery.
Figure 3: Quality–cost comparison on OpenCell.
FACET configuration
ε -adapter
Memory-conditioned semantic pathway
Backbone LoRA (base velocity vL )
Hierarchical head (velocity correction Δv )
IoUcont↑
FACET (full)
✓
✓
✓
✓
0.3328
Memory-conditioned + LoRA
✓
✓
✓
–
0.3236
Memory-conditioned + hier. head
✓
✓
–
✓
0.3052
Memory-conditioned
✓
✓
–
–
0.2826
ε -adaptation-only
✓
–
–
–
0.2581
Table 5: Controlled FACET configuration ablation on the HPA benchmark. Each configuration is evaluated on the same held-out proteins using identical cellular context, sampler, and preprocessing. Checkmarks indicate active pathways. “hier.” represents “hierarchical”.
Configuration
OpenCell adaptation
HPA module initialization
OpenCell-fit memory
IoUcont↑
FACET, HPA-adapted
–
✓
–
0.4222
FACET, OpenCell-adapted, cold start
✓
–
✓
0.4244
Warm start, inherited HPA memory
✓
✓
–
0.4155
Warm start, refitted OpenCell memory
✓
✓
✓
0.4233
Table 6: HPA-module transfer and OpenCell adaptation ablation. A checkmark indicates that the corresponding component is active. Note that, for all the configurations, the ε -predictor backbone, VAE encoder/decoder, and protein LM are adopted from the OpenCell checkpoint by Zheng and Huang (2025) .
Configuration
Learned memory
Memory replaced at inference
IoUcont↑
FACET, unmodified memory
✓
–
0.3328
FACET, uniform prototype mean
–
✓
0.2820
FACET, zero memory
–
✓
0.3014
Table 7: Ablation of inference-time memory on HPA. A checkmark indicates that the corresponding memory configuration is active.
Sampler
NFE
FID ↓
IoUcont↑
SDE
10
41.88
0.4301
SDE
25
36.83
0.4244
SDE
50
36.18
0.4214
ODE (Heun’s)
25
36.60
0.4205
ODE (Heun’s)
50
36.17
0.4187
Table 8: OpenCell sampling results across solvers and network function evaluations (NFE).
Method
ECE ↓
MCE ↓
Brier ↓
CELL-Diff
0.21501
0.24886
0.23372
FACET
0.05442
0.14850
0.19915
Table 9: Predictive calibration on HPA. One temperature per model, fit on a protein-disjoint validation split and held fixed on test. Foreground probabilities come from 16 predictive samples per input.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Semantic memory vocabulary
K=9 fixed UniProt localization categories
Residual representation
128 dimensions
Interpolant strength
κ=0.5
Correction head
2 residual blocks
LoRA adaptation
rank 16
Inference budget
25 NFE
Appendix
Table 10: Selected implementation settings for the FACET instantiation used in the reported experiments; component-level ablations are reported separately.
Figure 4: FACET inference overview. The protein sequence directly conditions the pretrained ε -predictor and queries the semantic memory to obtain mS , while a sequence-conditioned residual refines it into zS . The memory-conditioned ε -adapter forms the base velocity vL , and context-first hierarchical modulation produces the correction Δv . Their composition vΘ=vL+g(τ)Δv is iteratively integrated through stochastic latent transport, yielding the final latent estimate for decoding by the frozen VAE. Training-only paths are omitted; see Algorithm 2 for details.
Module
Parameters
Deployed
ε -adapter
1,318,533
✓
Prototype prior pη
923,145
✓
Hierarchical correction head
4,918,340
✓
LoRA (rank 16)
2,406,400
✓
Continuous prior pψ
262,785
✓
Total deployed
9,829,203
–
Appendix
Table 11: Parameter counts for the trainable FACET modules. Deployed modules are used during generation; posterior modules are used only during training.
Figure 5: Cellular-context ablation. The protein sequence is held fixed while the cellular-context input is removed or modified. The examples illustrate context-conditioned changes in the generated spatial phenotype; they are qualitative evidence rather than a causal ablation beyond the quantitative comparison reported in the main text.
Figure 6: Same-protein generation across HPA cellular contexts. The sequence input is fixed across contexts, allowing context-conditioned changes in the generated spatial phenotype to be visualized.
Figure 7: Additional HPA generations across multiple proteins and cellular contexts. The examples provide broader coverage of the conditional-generation setting beyond the main-text qualitative comparison.
Figure 8: Additional qualitative results on OpenCell. The figure uses the same method names and visualization conventions as the main OpenCell results and complements the quantitative comparison. Here “CELL-Diff (fine-tuned)” represents the CELL-Diff checkpoint fine-tuned on OpenCell data, the same one reported in the main text and the source ( Zheng and Huang, 2025 ) .
Figure 9: Additional qualitative results on HPA. CELL-E2 uses nucleus-only context, whereas CELL-Diff and FACET use the full three-channel context. These examples extend the main-text qualitative comparison using the same evaluation setting, method names, and visualization conventions.