pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows
Authors: Tong Chen, Maximilian Holsman, Lin Zhao, Pranam Chatterjee
Organizations: Department of Computer and Information Science, University of Pennsylvania · Department of Computer Science, Duke University · Department of Bioengineering, University of Pennsylvania
Biomolecular therapeutics often start from known sequences and require targeted editing to improve multiple properties while satisfying hard biochemical and manufacturability constraints. However, existing generative methods do not jointly support multi-objective optimization, hard feasibility, and sequence editing in discrete, variable-length biological spaces. In this work, we introduce Pareto-Constrained Molecule Editing (pCoMole), a framework built on discrete flow matching that steers a pre-trained Edit Flow toward user-specified preferences while enforcing terminal feasibility. pCoMole defines a feasibility-gated terminal distribution using an augmented Tchebycheff utility and realizes the resulting preference tilt through a Doob-h transform of the underlying edit process. To make this construction practical, we approximate the required harmonic function using short Monte Carlo rollouts over candidate edits, yielding an efficient guided editor with provable preference consistency. We validate pCoMole by shrinking GFP while retaining fluorescence-related properties, shortening diverse Cas9 orthologs while preserving PAM specificity, and compressing peptide binders into short peptidomimetics that optimize seven drug-related properties under hard constraints. In wet lab testing, two 229-residue pCoMole-designed eGFP variants retained clear green fluorescence in BL21 cells after 10 deletions, with either one or two substitutions. Together, pCoMole enables constraint-aware, Pareto-aligned editing of biomolecular sequences in discrete, variable-length spaces.
Figures & tables
Figure 1: Overview of pCoMole and representative designs. (A) Starting from an intermediate sequence xt , pCoMole samples candidate edits from a pretrained Edit Flows velocity ut and uses short Monte Carlo rollouts to estimate terminal utility and feasibility. The resulting approximate Doob- h guidance reweights transitions into a tilted velocity u~t , steering generation toward feasible, high-utility terminal sequences. (B) In GFP miniaturization, pCoMole reduces sequence length while preserving the canonical β -barrel and improving predicted optical objectives. (C) In Cas9 shrinkage, pCoMole produces a shortened SpCas9 variant that satisfies the PAM constraint while maintaining high predicted Cas9-likeness and plausible structure. (D) In short peptidomimetic binder design, pCoMole shortens a known 7JVS binder while improving predicted properties. The binder is shown in yellow, the target in light blue, and dark blue highlights target motif regions. Insets zoom into the binding interface (see Figure S6 for enlarged views).
Edit Flow
Train Loss
Test Loss
Diversity
Validity Rate
UniRef S
594.39
602.47
0.68
—
UniRef L
3621.89
3573.07
0.45
—
GFP
147.61
436.65
0.53
0.91
Cas9
2430.42
2498.37
0.25
0.93
Peptidomimetic
36.26
38.40
0.80
0.70
Table 1: Edit Flow training and evaluation metrics across sequence domains.
Method
∣λexc−488∣ ( ↓ )
Brightness
RayGun
22.3594
9.4207
SCISOR
20.1279
19.0333
pCoMole (UniRef S Edit Flow)
5.8213
43.5116
pCoMole (GFP Edit Flow)
8.4372
49.0121
Table 2: GFP design comparison between pCoMole, RayGun, and SCISOR under a fixed length constraint. For each method, we generate 100 candidates and report average metrics. All methods enforce a fixed terminal length of 213 residues and achieve a validity rate near 100%.
Task
Method
Low Tier (s)
Medium Tier (s)
High Tier (s)
GFP
RayGun
N.A.
N.A.
N.A.
SCISOR
14
613
N.A.
pCoMole
266
333
665
Cas9
RayGun
N.A.
N.A.
N.A.
SCISOR
6.86
76.61
240.28
pCoMole
131.69
129.15
138.96
Table 3: Compute-normalized time to first predicted-usable sequence. We report the time required to obtain a sequence satisfying each task-specific quality tier under a 3600-second trial budget. Lower is better. “N.A.” indicates that no valid sequence was found within the budget.
Objective Guidance ∣λexc−488∣ Brightness
∣λexc−488∣ ( ↓ )
Brightness
Length
Validity Rate
✓✓
3.7557
54.0159
214
1.0
✓×
2.9810
31.2600
213
1.0
×✓
12.9500
62.1200
214
1.0
××
32.8200
25.0200
213
1.0
Table 4: Objective guidance ablation for pCoMole on GFP design. For each objective guidance setting, we report the average metrics over 100 pCoMole-shrunk GFP sequences. Length reduction guidance is applied in all settings.
Figure 2: Wet-lab fluorescence validation of pCoMole-shortened eGFP designs. 488-nm fluorescence images of the 239-residue in-house eGFP construct, a manually selected 228-residue negative control, and two 229-residue pCoMole designs. pCoMole_eGFP_1 contains 10 deletions and one substitution; pCoMole_eGFP_2 contains 10 deletions and two substitutions.
Input Cas9
Length Change
Cas9 Likelihood
PAM Distr. CE
PAM Match Rate
PAM matched
SpCas9
1368→1211(−157)
1.00→0.97
0.97→0.90
1.00
NGG
St1Cas9
1121→1023(−98)
1.00→0.98
0.91→0.89
1.00
NNANAA
St3Cas9
1388→1301(−87)
1.00→0.95
0.94→0.90
1.00
NGGNG
GeoCas9
1087→1035(−52)
0.96→0.95
0.93→0.89
1.00
NNNNCNAA
Table 5: pCoMole shrinks common Cas9 sequences while maintaining PAM specificity. “Cas9 Likelihood” is the score of a pre-trained Cas9 classifier. “PAM Distribution CE” is a cross-entropy between the predicted PAM distribution of the input and that of the generated sequence. “PAM Match Rate” is the fraction of generated sequences whose predicted PAM exactly matches the input. “PAM matched’ is the predicted PAM which pCoMole maintained during shrinkage.
Target
Ligand
Length
Non-Toxicity
Solubility
Permeability
Half-life (h)
Affinity
Motif Score
Specificity
PTH1R
Teriparatide
259
0.8890
0.6905
0.2793
2.8301
0.9148
0.2696
0.9834
pCoMole
217
0.8729
0.7268
0.2959
2.9488
0.9358
0.3593
0.9706
p53
p28
170
0.8533
0.6737
0.2627
2.2015
0.6369
0.5418
0.9873
pCoMole
133
0.8076
0.6309
0.3071
3.8517
0.6885
0.5014
0.9649
GLP-1R
Semaglutide
248
0.8872
0.7656
0.2583
1.8571
0.9467
0.4691
0.9827
pCoMole
186
0.8122
0.7240
0.2985
4.0277
0.9367
0.5212
0.9778
Table 6: pCoMole designs shorter peptidomimetics with improved predicted properties. For each target, we compare the original peptide drug binder with the average over 100 pCoMole-designed peptidomimetics. Length denotes the number of tokenized SMILES tokens. All pCoMole designs satisfy the specified task constraints. For GLP-1R, motif score and specificity are computed relative to Semaglutide and are therefore reported only for Semaglutide and its pCoMole designs.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
PDB
Constraints
Length
Non-Toxicity
Solubility
Permeability
Half-life (h)
Affinity
Motif Score
Specificity
RPM
RLEN
3AMA
Baseline (Known Binder)
141
0.901
0.854
0.2528
1.4508
0.6731
0.9171
0.9941
—
—
PM+LEN
129
0.8085
0.4991
0.3049
4.8599
0.6946
0.9798
0.9429
1.00
1.00
PM only
112
0.7444
0.5769
0.3537
14.4684
0.6666
0.8521
0.9482
1.00
0.27
LEN only
131
0.8261
0.486
0.3081
4.9182
0.6947
0.9782
0.9443
1.00
1.00
None
21
0.5104
0.8422
0.6186
68.4154
0.4696
0.4247
0.9613
0.10
0.02
5JHF
Baseline (Known Binder)
95
0.8029
0.8744
0.2431
1.5936
0.5457
0.8484
0.9926
—
—
Appendix
Table S1: Constraint ablation for peptidomimetic binder design on 3AMA and 5JHF. For each target, we report the property scores of the known binder (“Baseline”) and the average scores over 100 pCoMole-generated sequences under different constraint settings. Length is the number of tokenized SMILES tokens. PM denotes the hard peptidomimetic constraint and LEN denotes the soft length constraint. RPM and RLEN are the fractions (out of 100) of generated sequences that satisfy the PM and LEN constraints, respectively. Please refer to Section D for details.
Table S2: Pareto coverage comparison on PDB 4O56 under different scalarizations and weight vectors. We evaluate three scalarization choices (ATC, Tchebycheff, and linear weighted sum) for multi-objective peptidomimetic binder design optimizing solubility, permeability, and binding affinity. For each method, we run pCoMole with three sets of objective weight vectors and report the average properties of 100 generated sequences. The rightmost column reports each method’s empirical Pareto coverage against a reference front constructed by pooling solutions across all weight settings and methods (higher is better).
Model
Length Diff.
Cas9 Likeli.
PAM Distr. CE
PAM Match %
Base EF ( s=1 )
1121→1119
1.00→1.00
0.91→0.89
36
Base EF ( s=1000 )
1121→1004
1.00→0.70
0.91→0.78
0
pCoMole ( s=1 )
1121→1118
1.00→1.00
0.91→0.94
100
pCoMole ( s=1000 )
1121→1016
1.00→0.95
0.91→0.89
100
Appendix
Table S3: Deletion-rate scaling ablation for pCoMole on St1Cas9 edits. In all cases, the corresponding model was used to generate 50 sequences. “Base EF” denotes sampling from the Edit Flows directly without any guidance. “s” denotes the deletion rate scaling term.
pCoMole Config.
Length Diff.
Cas9 Likeli.
PAM Distr. CE
PAM Match %
pCoMole w/ no ins.
1121→1021
1.00→0.98
0.91→0.89
100
pCoMole w/ ins.
1121→1049
1.00→0.97
0.91→0.90
100
Appendix
Table S4: Ablation of allowing pCoMole edit steps to make insertions. In both cases, pCoMole was used to generate 50 shrunk variants of St1Cas9.
PDB
#-steps
Length
Non-Toxicity
Solubility
Permeability
Half-life (h)
Affinity
Motif Score
Specificity
Time (s)
2KXQ
0
130
0.5343
0.8607
0.2882
2.2298
0.6373
0.9528
0.8556
—
10
116
0.8017
0.787
0.3000
3.5262
0.7565
0.9719
0.7786
69.73
20
111
0.7683
0.7279
0.3113
4.4267
0.8055
0.9690
0.7882
150.94
30
108
0.7626
0.7036
0.3180
5.3392
0.8106
0.9705
0.7837
257.37
50
99
0.7541
0.6664
0.3330
6.7532
0.8037
0.9699
0.7928
530.83
8EYA
0
114
0.7220
0.8249
0.2410
4.2260
0.6285
0.9392
0.9967
—
Appendix
Table S5: Sampling-step ablation for peptidomimetic binder design on 2KXQ and 8EYA. For each target, the row with #-steps =0 reports the known binder (baseline), and the remaining rows report averages over 100 pCoMole-generated sequences under different numbers of sampling steps. Length is the number of tokenized SMILES tokens. We also report the average wall-clock time (seconds) to generate one sequence for each step setting.
PDB
#-rollouts
Length
Non-Toxicity
Solubility
Permeability
Half-life (h)
Affinity
Motif Score
Specificity
Time (s)
2QSC
0
111
0.8861
0.7156
0.2874
1.7210
0.6775
0.9323
0.9621
—
5
87
0.7084
0.6840
0.2940
5.9379
0.7745
0.9371
0.9470
78.79
10
86
0.7172
0.6858
0.2974
6.2827
0.7809
0.9287
0.9477
124.57
20
87
0.7447
0.6938
0.3063
6.7057
0.7826
0.9335
0.9464
216.92
30
84
0.7175
0.6967
0.3042
7.1078
0.7756
0.9294
0.9462
306.83
7ZPY
0
84
0.4621
0.7744
0.2787
2.5357
0.5218
0.3083
0.9954
—
Appendix
Table S6: Rollout ablation for peptidomimetic binder design on 2QSC and 7ZPY. For each target, the row with #-rollouts =0 reports the known binder (baseline), and the remaining rows report averages over 100 pCoMole-generated sequences under different numbers of rollouts. Length is the number of tokenized SMILES tokens. We also report the average wall-clock time (seconds) to generate one sequence for each rollout setting.
PDB
#-candidates
Length
Non-Toxicity
Solubility
Permeability
Half-life (h)
Affinity
Motif Score
Specificity
Time (s)
2O9V
0
65
0.0826
0.8305
0.2977
2.7590
0.4128
0.9671
0.8806
—
10
47
0.6998
0.7226
0.4626
16.5440
0.6143
0.9692
0.8327
41.42
20
45
0.7330
0.7345
0.4633
17.5760
0.6241
0.9646
0.8309
76.76
30
46
0.7511
0.7323
0.4574
19.6885
0.6278
0.9658
0.8339
58.32
50
47
0.7532
0.7379
0.4541
20.1591
0.6242
0.9683
0.8312
163.55
8PFT
0
143
0.6156
0.7538
0.185
6.1210
0.5750
0.6579
0.9502
—
Appendix
Table S7: Number of candidates ablation for peptidomimetic binder design on 2O9V and 8PFT. For each target, the row with #-candidates =0 reports the known binder (baseline), and the remaining rows report averages over 100 pCoMole-generated sequences under different numbers of candidates. Length is the number of tokenized SMILES tokens. We also report the average wall-clock time (seconds) to generate one sequence for each candidate-count setting.
Ortholog
PAM guidance
Protein2PAM exact match
Protein2PAM PWM sim.
CICERO exact match
CICERO PWM sim.
St3Cas9
Enabled
100 [87, 100]
0.896
88 [70, 96]
0.837
St3Cas9
Disabled
16 [6, 35]
0.634
4 [1, 20]
0.758
GeoCas9
Enabled
100 [87, 100]
0.792
72 [52, 86]
0.766
GeoCas9
Disabled
88 [70, 96]
0.392
20 [9, 39]
0.619
Appendix
Table S8: Held-out CICERO evaluation of pCoMole-generated Cas9 designs. We evaluated 25 designs per ortholog and condition. The guidance-disabled control removes the PAM objective and hard PAM constraint while keeping the remaining pCoMole settings fixed. Exact-match values are percentages with Wilson 95% confidence intervals; PWM similarity compares the full predicted PAM distributions.
Method
Length Diff.
Avg. Cas9 Likelihood
PAM Match %
pCoMole w/ Cas9 Edit Flows
984 ⟶ 934
0.95
100
pCoMole w/ UniRef Edit Flows
984 ⟶ 934
0.98
100
SCISOR
984 ⟶ 934
0.82
35
RayGun
984 ⟶ 934
0.0
0
Appendix
Table S9: Comparison of pCoMole to SCISOR and RayGun at shrinking CjCas9. All three methods were used to shrink WT CjCas9 to a sequence length of 934.
PDB
Length Change
Δ Non-Toxicity
Δ Solubility
Δ Permeability
Δ Half-life (h)
Δ Affinity
Δ Motif Score
Δ Specificity
1AYC
59 → 38
-0.1288
+0.0163
+0.1436
+12.8558
+0.1001
+0.0268
-0.0386
1B8Q
51 → 33
-0.2324
+0.0456
+0.1512
+33.3370
-0.0643
+0.0024
-0.0389
1DDV
41 → 28
+0.3973
+0.0293
+0.1945
+31.1299
+0.2025
-0.0214
+0.0098
1E6I
43 → 29
-0.1412
+0.1025
+0.1748
+23.5607
+0.1515
+0.0547
-0.0537
2LTV
85 → 64
+0.3128
-0.0909
+0.0584
+9.6710
+0.1178
+0.1034
-0.0978
2Q8Y
72 → 49
-0.0639
-0.0591
+0.0971
+18.0023
+0.2046
+0.0122
-0.0437
Appendix
Table S10: pCoMole designs short peptidomimetic binders for 15 PDB targets with known peptide binders, achieving substantial length reduction and improved predicted properties. For each target, we report the average scores over 100 designed peptidomimetics and the property score change ( Δ ) relative to the pre-existing binder. Length is the number of SMILES tokens after tokenization. All designed sequences satisfy the specified constraints, and positive Δ indicates improvement.
Figure S1: Distribution of average pLDDT scores across a set of 1000 sequences generated by the UniRef S Edit Flow, compared to the 1000 true Uniref sequences used as inputs. 1000 UniRef sequences with length ≤350 were sampled and used as x0 for these generations.
Figure S2: Objective guidance controls optical trade-offs while preserving GFP structure under shrinkage. Representative pCoMole-designed GFP variants generated under four objective-guidance settings: (A) brightness, excitation alignment, and length reduction (B) brightness and length reduction (C) excitation alignment and length reduction (D) length reduction only. We present their secondary structures together with AlphaFold3-predicted pTM scores, sequence length, predicted brightness, and excitation alignment error.
Figure S3: Complex structures of known peptide binders and pCoMole-designed peptidomimetic binders with three target PDB proteins (A) 1AYC (B) 1B8Q (C) 1DDV. Binders are shown in yellow, target proteins in light blue, and the dark-blue surface highlights the target motifs used by pCoMole to enforce motif-specific binding during design. Insets zoom into the binding interface to illustrate key contacts. Predicted property scores for each binder are reported.
Figure S4: Complex structures of known peptide binders and pCoMole-designed peptidomimetic binders with three target PDB proteins (D) 1E6I (E) 2LTV (F) 2Q8Y. Binders are shown in yellow, target proteins in light blue, and the dark-blue surface highlights the target motifs used by pCoMole to enforce motif-specific binding during design. Insets zoom into the binding interface to illustrate key contacts. Predicted property scores for each binder are reported.
Figure S5: Complex structures of known peptide binders and pCoMole-designed peptidomimetic binders with three target PDB proteins (G) 4GNE (H) 5AZ8 (I) 6MLC. Binders are shown in yellow, target proteins in light blue, and the dark-blue surface highlights the target motifs used by pCoMole to enforce motif-specific binding during design. Insets zoom into the binding interface to illustrate key contacts. Predicted property scores for each binder are reported.
Figure S6: Complex structures of known peptide binders and pCoMole-designed peptidomimetic binders with three target PDB proteins (J) 7JVS (K) 7LUL (L) 8CN1. Binders are shown in yellow, target proteins in light blue, and the dark-blue surface highlights the target motifs used by pCoMole to enforce motif-specific binding during design. Insets zoom into the binding interface to illustrate key contacts. Predicted property scores for each binder are reported.
Figure S7: pCoMole effectively shrinks a diverse set of Cas9 orthologs. Representative pCoMole-designed Cas9 sequences generated by shrinking four well-characterized Cas9 orthologs: (A) SpCas9, (B) St1Cas9 (C) St3Cas9 (D) GeoCas9. We present their secondary structures together with AlphaFold3-predicted pTM scores, sequence length, Cas9-likelihood score, and predicted PAM.
Small-molecule drug discovery requires simultaneous optimization of numerous properties of candidate molecules. These properties can be investigated through the analysis of high-dimensional biological signatures, such as cell morphology and transcriptomic perturbations, which provide a rich perspective on the underlying biological mechanisms. However, existing generative methods, which use those signatures for optimization, fail to meet two key requirements: providing precise guidance toward desired phenotypic signatures while maintaining structural proximity to a known hit. We introduce PhAME (Phenotype-Aware Molecular Editing), a latent diffusion framework that overcomes this challenge by recasting molecular optimization as editing in the latent space of a pretrained graph-based VAE. Our central contribution is a compositional classifier-free guidance scheme with two independent scales, one for the phenotype-conditioning and one for similarity to the seed structure, allowing practitioners to control the tradeoff between these two objectives. Empirical evaluations across diverse benchmarks, including docking score optimization and multimodal phenotypic generation, demonstrate that PhAME achieves state-of-the-art results while maintaining high chemical validity and novelty.
Łukasz Janisiów, Sebastian Musiał, Bartosz Zieliński +2
Faculty of Mathematics and Computer Science, Jagiellonian University · Doctoral School of Exact and Natural Sciences, Jagiellonian University · Jagiellonian Center for Artificial Intelligence, Jagiellonian University +2
Conditional molecular optimization aims to edit a molecule to realize a specified property shift. In practice, structurally similar molecule data is scarce, while decisions are inherently action-level: at each step, the system must select one local structural edit from a candidate set that is strictly filtered by chemical feasibility rules. This level mismatch between supervision and decision makes oracle-in-the-loop search unstable in molecular optimization. Regressing on property differences between molecule pairs improves data efficiency but relies on oracle-in-the-loop search, entangling transformation effects with global context and providing limited guidance for selecting the next feasible edit, often resorting to oracle-in-the-loop search. For this reason, we propose a response-oriented discrete edit optimization approach comprising two tightly coupled components: a single-step molecular edit response predictor (SMER) and a multi-step planner that composes local predictions into optimization trajectories via guided tree search (SMER-Opt). The approach learns a directional evaluation model over edit actions to support constraint-aware planning. It mines weakly related molecule pairs and decomposes their structural differences into minimal edit units, turning endpoint property annotations into process-level supervision and yielding reusable, transferable action primitives. A directional edit evaluator then scores feasible candidate edits by their likelihood of moving the molecule toward the desired property change, substantially reducing dependence on external evaluator queries at decision time. Code is available at https://anonymous.4open.science/r/SMER.
Haojie Rao, Kun Li, Yida Xiong +5
School of Computer Science, Wuhan University, Wuhan, China · School of Computer Science, Wuhan University Wuhan, China · Department of Data Science and Artificial Intelligence, Monash University, Victoria, Australia +5
Designing functional biological sequences requires navigating vast discrete spaces under strict evolutionary and biophysical constraints. Discrete Flow Matching (DFM) offers a generative framework over such spaces, but existing approaches rely on biologically uninformative couplings and offer limited flexibility for variable-length sequence generation and fine-grained control. We propose a structured coupling that encodes domain-specific preferences among sequence elements, biasing the source distribution toward plausible regions without modifying the flow objective or training procedure. Building on this, we introduce a latent edit-based rate parameterization that models variable-length generation via edit operations conditioned on a shared global latent, akin to a latent variable model, while remaining tractable. We further introduce a latent classifier-free guidance mechanism that steers generation coherently in continuous latent space, along with Dirichlet-prior temperature scaling for test-time control over edit operations. Our method achieves state-of-the-art performance across diverse biological sequence tasks, including density estimation, unconditional and conditional DNA sequence generation, and peptide sequence generation.
Yogesh Verma, Dani Korpela, Harri Lähdesmäki +1
Aalto University, Finland · Aalto University and YaiYai Ltd