Fluorescence microscopy reveals where proteins localize, but only a limited number of proteins can be imaged in the same cell; generating these images from amino-acid sequence and the cell's morphological context enables in silico localization of unimaged proteins. The two conditions, however, play asymmetric roles: morphological context is spatially aligned with the target, whereas sequence is non-spatial and must specify protein-dependent localization within it, with recurring coarse patterns shared across proteins and finer protein-specific variation. Existing generators condition on both jointly, without separating what each explains. We introduce FACET (Factorized Asymmetric Conditioning for Efficient Transport), a probabilistic generative framework that encodes this structure as an explicit inductive bias: sequence semantics are learned from what context leaves unexplained, coarse localization regularities are shared across proteins through a semantic memory, and protein-specific variation is a bounded residual around them. A variance-preserving state projection further lets FACET perform continuous stochastic transport through a pretrained diffusion predictor with minimal parameter overhead. On held-out proteins, FACET improves spatial overlap by 34.3% on the Human Protein Atlas and 14.0% on OpenCell over a backbone-matched baseline, and reduces FID by 27.2% and 46.5%, respectively, with 75% fewer network evaluations. It also substantially improves protein-association structure recovery and yields better-calibrated predictions, while detailed ablations show complementary contributions from its design choices. These results identify factorized asymmetric conditioning, rather than generator capacity alone, as a key lever for high-fidelity, efficient, and biologically meaningful cellular image synthesis.
Figures & tables
Figure 1: Probabilistic graphical model for FACET .
Figure 2: Representative held-out HPA examples: morphology context, target, and model outputs. CELL-E2 uses nucleus-only context; CELL-Diff and FACET use all three context channels.
Method
NFE ↓
MI ↑
MS-SSIM ↑
Haralick ↑
PSNR ↑
IoUcont↑
FID ↓
CELL-E2
–
0.4041
0.1370
0.5644
6.36
0.1755
166.4
CELL-Diff
100
0.5355
0.3471
0.5182
15.77
0.2478
45.60
FACET
25
0.5481
0.3651
0.5854
16.15
0.3328
33.20
Table 1: Comparison on the HPA test set. Higher is better for all metrics except NFE and FID. Best results are shown in bold.
Context
System
MI ↑
MS-SSIM ↑
Haralick ↑
PSNR ↑
IoUcont↑
FID ↓
Nucleus
CELL-Diff
0.3091
0.2777
0.4932
15.32
0.2047
51.10
FACET
0.2519
0.1712
0.5010
14.66
0.2768
52.12
Nucleus + ER
CELL-Diff
0.4562
0.3154
0.4898
15.52
0.2182
60.00
FACET
0.5120
0.3413
0.5726
15.96
0.3268
33.83
Nucleus + MT
CELL-Diff
0.5097
0.3404
0.5196
15.60
0.2435
47.60
FACET
0.5226
0.3524
0.5796
16.03
0.3266
33.24
Table 2: Performance across cellular-context compositions on HPA. All systems use the same proteins, sequence inputs, and sampling protocol per composition. Bold marks the better value within each composition.
Method
NFE ↓
MI ↑
MS-SSIM ↑
Haralick ↑
PSNR ↑
IoU cont ↑
FID ↓
CELL-E2
–
0.1649
0.1020
0.4542
11.31
0.1779
245.8
CELL-Diff
100
0.2732
0.4078
0.3663
13.20
0.3722
68.90
FACET (HPA-adapted)
25
0.2753
0.3653
0.5177
13.85
0.4222
43.33
FACET (OpenCell-adapted)
25
0.2998
0.3893
0.5199
13.75
0.4244
36.83
Table 3: Comparison on the OpenCell test set (nucleus-only context). All systems except CELL-E2 share the OpenCell-trained CELL-Diff backbone; HPA-adapted transfers the FACET modules from HPA without refitting, whereas OpenCell-adapted fits them on OpenCell. Best results are in bold.
System
STRING score threshold
Pairs (assoc./null)
Pearson Δ↑
passoc.vs.null
Baseline
0.85
47/2659
+0.1005
0.298
FACET
0.85
47/2659
+0.1875
1.3×10−3
Baseline
0.70
65/2641
+0.0954
0.170
FACET
0.70
65/2641
+0.1718
4.5×10−4
Baseline
0.50
112/2594
+0.0922
0.042
FACET
0.50
112/2594
+0.1369
2.5×10−4
Table 4: Spatially resolved protein-association pattern recovery under a localization-matched evaluation adapted from ( Sun et al., 2026 ) . The baseline denotes CELL-Diff. Pearson similarity is computed after removing the per-cell mean image. Δ measures the separation between STRING-associated and matched non-associated pairs; positive values indicate stronger recovery.
Figure 3: Quality–cost comparison on OpenCell.
FACET configuration
ε -adapter
Memory-conditioned semantic pathway
Backbone LoRA (base velocity vL )
Hierarchical head (velocity correction Δv )
IoUcont↑
FACET (full)
✓
✓
✓
✓
0.3328
Memory-conditioned + LoRA
✓
✓
✓
–
0.3236
Memory-conditioned + hier. head
✓
✓
–
✓
0.3052
Memory-conditioned
✓
✓
–
–
0.2826
ε -adaptation-only
✓
–
–
–
0.2581
Table 5: Controlled FACET configuration ablation on the HPA benchmark. Each configuration is evaluated on the same held-out proteins using identical cellular context, sampler, and preprocessing. Checkmarks indicate active pathways. “hier.” represents “hierarchical”.
Configuration
OpenCell adaptation
HPA module initialization
OpenCell-fit memory
IoUcont↑
FACET, HPA-adapted
–
✓
–
0.4222
FACET, OpenCell-adapted, cold start
✓
–
✓
0.4244
Warm start, inherited HPA memory
✓
✓
–
0.4155
Warm start, refitted OpenCell memory
✓
✓
✓
0.4233
Table 6: HPA-module transfer and OpenCell adaptation ablation. A checkmark indicates that the corresponding component is active. Note that, for all the configurations, the ε -predictor backbone, VAE encoder/decoder, and protein LM are adopted from the OpenCell checkpoint by Zheng and Huang (2025) .
Configuration
Learned memory
Memory replaced at inference
IoUcont↑
FACET, unmodified memory
✓
–
0.3328
FACET, uniform prototype mean
–
✓
0.2820
FACET, zero memory
–
✓
0.3014
Table 7: Ablation of inference-time memory on HPA. A checkmark indicates that the corresponding memory configuration is active.
Sampler
NFE
FID ↓
IoUcont↑
SDE
10
41.88
0.4301
SDE
25
36.83
0.4244
SDE
50
36.18
0.4214
ODE (Heun’s)
25
36.60
0.4205
ODE (Heun’s)
50
36.17
0.4187
Table 8: OpenCell sampling results across solvers and network function evaluations (NFE).
Method
ECE ↓
MCE ↓
Brier ↓
CELL-Diff
0.21501
0.24886
0.23372
FACET
0.05442
0.14850
0.19915
Table 9: Predictive calibration on HPA. One temperature per model, fit on a protein-disjoint validation split and held fixed on test. Foreground probabilities come from 16 predictive samples per input.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Semantic memory vocabulary
K=9 fixed UniProt localization categories
Residual representation
128 dimensions
Interpolant strength
κ=0.5
Correction head
2 residual blocks
LoRA adaptation
rank 16
Inference budget
25 NFE
Appendix
Table 10: Selected implementation settings for the FACET instantiation used in the reported experiments; component-level ablations are reported separately.
Figure 4: FACET inference overview. The protein sequence directly conditions the pretrained ε -predictor and queries the semantic memory to obtain mS , while a sequence-conditioned residual refines it into zS . The memory-conditioned ε -adapter forms the base velocity vL , and context-first hierarchical modulation produces the correction Δv . Their composition vΘ=vL+g(τ)Δv is iteratively integrated through stochastic latent transport, yielding the final latent estimate for decoding by the frozen VAE. Training-only paths are omitted; see Algorithm 2 for details.
Module
Parameters
Deployed
ε -adapter
1,318,533
✓
Prototype prior pη
923,145
✓
Hierarchical correction head
4,918,340
✓
LoRA (rank 16)
2,406,400
✓
Continuous prior pψ
262,785
✓
Total deployed
9,829,203
–
Appendix
Table 11: Parameter counts for the trainable FACET modules. Deployed modules are used during generation; posterior modules are used only during training.
Figure 5: Cellular-context ablation. The protein sequence is held fixed while the cellular-context input is removed or modified. The examples illustrate context-conditioned changes in the generated spatial phenotype; they are qualitative evidence rather than a causal ablation beyond the quantitative comparison reported in the main text.
Figure 6: Same-protein generation across HPA cellular contexts. The sequence input is fixed across contexts, allowing context-conditioned changes in the generated spatial phenotype to be visualized.
Figure 7: Additional HPA generations across multiple proteins and cellular contexts. The examples provide broader coverage of the conditional-generation setting beyond the main-text qualitative comparison.
Figure 8: Additional qualitative results on OpenCell. The figure uses the same method names and visualization conventions as the main OpenCell results and complements the quantitative comparison. Here “CELL-Diff (fine-tuned)” represents the CELL-Diff checkpoint fine-tuned on OpenCell data, the same one reported in the main text and the source ( Zheng and Huang, 2025 ) .
Figure 9: Additional qualitative results on HPA. CELL-E2 uses nucleus-only context, whereas CELL-Diff and FACET use the full three-channel context. These examples extend the main-text qualitative comparison using the same evaluation setting, method names, and visualization conventions.
De novo protein generation has transformative potential in therapeutic design, enzyme engineering, and synthetic biology. While diffusion-based and flow matching approaches have achieved progress, they typically operate at single resolution and lack mechanisms for incorporating functional constraints. We introduce ProHiFlo, a hierarchical flow matching framework with three innovations: (1) coarse-to-fine generation that models backbone geometry before refining to all-atom coordinates, reducing computational cost while maintaining accuracy; (2) functional guidance leveraging pretrained predictors to steer generation toward desired properties without retraining; (3) adaptive SE(3)-equivariant architecture for efficient multi-scale processing. Experiments on unconditional generation, motif scaffolding, and functional design demonstrate state-ofthe-art performance while requiring 4 fewer sampling steps. On enzyme active site scaffolding, ProHiFlo achieves 58.9% success rate compared to 41.2% for RFDiffusion.
Chuanzhen Wang, Meade Cleti, Pete Jano
3Tongji University · 1Arizona State University · University of Wisconsin-Madison
Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has strong low-rank structure. We present Asymmetric Flow Modeling (AsymFlow), a rank-asymmetric velocity parameterization that restricts noise prediction to a low-rank subspace while keeping data prediction full-dimensional. From this asymmetric prediction, AsymFlow analytically recovers the full-dimensional velocity without changing the network architecture or training/sampling procedures. On ImageNet 256×256, AsymFlow achieves a leading 1.57 FID, outperforming prior DiT/JiT-like pixel diffusion models by a large margin. AsymFlow also provides the first-ever route for finetuning pretrained latent flow models into pixel-space models: aligning the low-rank pixel subspace to the latent space gives a seamless initialization that preserves the latent model's high-level semantics and structure, so finetuning mainly improves low-level mismatches rather than relearning pixel generation. We show that the pixel AsymFlow model finetuned from FLUX.2 klein 9B establishes a new state of the art for pixel-space text-to-image generation, beating its latent base on HPSv3, DPG-Bench, and GenEval while qualitatively showing substantially improved visual realism.
Computational enzyme design requires generating proteins that scaffold catalytic residues and ligands, a task that demands both geometric accuracy and structural diversity from the underlying generative model. Current all-atom generators inherit expensive architectures from structure prediction, leading to high training costs and limited sample diversity. We argue that much of this complexity is unnecessary for generators, which condition on sparse geometric constraints rather than rich co-evolutionary signals. Emyx is a 140M-parameter conditional flow matching model that concentrates capacity within standard transformer blocks, replacing heavy embedding stacks with lightweight conditional representations and sparse connectivity. We additionally derive an exact reparametrisation of the flow matching interpolant into the EDM noise-level framework, bridging flow matching training efficiency with state-of-the-art sampling methods designed for diffusion models without retraining. Despite being the smallest model, Emyx outperforms both Proteína-Complexa and RFdiffusion3 against the AME enzyme design benchmark across success rate under strict evaluation requiring both global fold recovery and catalytic geometry accuracy, structural novelty, scaffold diversity, and geometric validity, while training in just 682 GPU-hours, roughly 4× less than RFdiffusion3.
Nicholas J. Williams, Ward Haddadin, Matteo P. Ferla +6