Bridging the Evidence-to-Execution Gap:A Reflective Agent for Multi-Objective Peptide Design
Organizations: School of Agriculture and Biology, Shanghai Jiao Tong University · School of Computer Science, Shanghai Jiao Tong University
Abstract
Large language models (LLMs) can reason over scientific literature to devise design strategies, yet fail to reliably implement them for biological sequences. While protein generative models learn sequence patterns, they lack the capacity to incorporate literature evidence for multi-step reflective reasoning, forming an evidence-to-execution gap between scientific reasoning and sequence manipulation. We present EASER (Evidence-Aware Sequence Engineering with Reflection), a reflective agent bridging reasoning and sequence generation via a learned property interface of offline-trained, fixed low-rank matrices. The agent steers a diffusion generator by combining these matrices, proposing intervention hypotheses (anchors, editable positions, control coefficients) grounded in retrieved evidence, sequence context and past results. A Probe-and-Steer mechanism validates interventions and allocates samples according to predicted property responses, with outcome reflection informing subsequent decisions. Evaluated on multi-objective antimicrobial peptide design (optimizing activity, non-hemolysis and non-toxicity), explicit hypothesis formulation delivers better multi-objective performance than direct action generation under identical decision conditions. Ablation studies verify the importance of evidence retrieval, episodic history, reflection and Probe-and-Steer. Over six repeated trials, EASER obtains the highest mean hypervolume and lowest mean IGD+ on screened candidates compared with competing baselines. Our work demonstrates how an executable property interface and iterative feedback link scientific reasoning to targeted peptide sequence generation.
Figures & tables
| Method | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|
| AMPGAN | |||
| AMPGPT2 | |||
| AMPGen | |||
| AMPDesigner | |||
| AMP-adapted DPLM | |||
| EASER |
| Method | Seeds | Passed | HV | IGD+ |
|---|---|---|---|---|
| AMPGAN | ||||
| AMPGPT2 | ||||
| AMPGen | ||||
| AMPDesigner | ||||
| AMP-adapted DPLM | ||||
| EASER |
| Method | Exact HV | IGD+ |
|---|---|---|
| Fixed coefficients | ||
| Random actions | ||
| LinUCB | ||
| NSGA-II | ||
| EASER |
| Policy | Mean | Screened candidates |
|---|---|---|
| Direct action generation | ||
| Explicit hypothesis formulation | 0.002447 | 261 |
| Condition | HV | IGD+ | Passed |
|---|---|---|---|
| EASER | |||
| Without reflection feedback | |||
| Without round-history details |
| Method | HV | IGD+ | Pass |
|---|---|---|---|
| EASER | |||
| Without memory and anchor reuse | |||
| Without retrieval | |||
| Without Probe-and-Steer | |||
| All matrices randomized |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Target | Predictor | Training data and method |
|---|---|---|
| Internal feedback | ||
| Activity | LLAMP ( Bae et al., 2025 ) | Released MIC-trained model combining peptide ESM-2 representations with bacterial genomic features. |
| Hemolysis | Calibrated ConsAMPHemo ( Xie et al., 2025 ) | Released HC50 regression model; we fit a gradient-boosted hemolysis classifier using its score and physicochemical features. |
| Toxicity | ToxinPred3 ( Rathore et al., 2024 ) | Released toxic/non-toxic sequence classifier; the native hybrid score supplies toxicity feedback. |
| External evaluation | ||
| Activity | MIC regression model | Augmented experimental MIC records from four bacterial species; ESM-2 and sequence features feed a neural regressor. Evaluated for E. coli . |
| Method | Exact HV | IGD+ |
|---|---|---|
| Search-policy baselines | ||
| Fixed coefficients | ||
| Random actions | ||
| LinUCB | ||
| NSGA-II | ||
| External generators | ||
| Use | Source and processing |
|---|---|
| Generator adaptation | Reviewed UniProtKB/Swiss-Prot ( The UniProt Consortium, 2025 ) short proteins and mature peptides, followed by a curated AMP collection. Sequences are deduplicated and restricted to 7–50 residues; the Swiss-Prot pool is split by sequence clusters. |
| Activity preferences | Experimental MIC records from AMPBench-MT ( Zhou et al., 2026b ) , including DBAASP, GRAMPA, CAMPR4, and DRAMP sources ( Pirtskhalava et al., 2021 ; Witten and Witten, 2019 ; Gawde et al., 2023 ; Shi et al., 2022 ) . The cleaned E. coli subset supplies physicochemically matched preference pairs. |
| Safety preferences | A combined experimental HC50 collection and a curated binary toxicity collection. Sequences are cleaned, filtered, and paired by the corresponding safety labels and physicochemical features. |
| Activity evaluator | Integrated experimental MIC data, supplemented with GRAMPA ( Witten and Witten, 2019 ) , AMP-Diffusion data ( Torres et al., 2025 ) , and published in-vitro measurements. Sequence, species, and MIC records are consolidated for regression. |
| AMP reference | AMP-Diffusion ( Torres et al., 2025 ) , AMPBench-MT ( Zhou et al., 2026b ) , and curated peptide activity collections are merged and deduplicated. The 27,733 sequences of length 15–35 provide the physicochemical KDE reference. |
| Agent knowledge | Curated property evidence, amino-acid guidance, motif-risk cards, and intervention strategies are organized into retrievable entries. |
| Setting | Stage 1 Short peptides | Stage 2 AMP-oriented |
|---|---|---|
| Train / validation records | 7,453 / 848 | 8,725 / 1,255 |
| LoRA rank / scaling alpha | 8 / 16 | 4 / 8 |
| LoRA dropout | 0.05 | 0.10 |
| Target projections | Query, value | Query, value |
| Learning rate | ||
| Warmup steps | 300 | 150 |
| Property | Preference pairs | Selected step | Parameters | ||
|---|---|---|---|---|---|
| Train | Valid. | Test | |||
| Activity | |||||
| Low hemolysis | |||||
| Low toxicity | |||||
| Activity score | Non-hemolysis score | Non-toxicity score | |
|---|---|---|---|
| Active matrix | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|
| Activity | |||
| Low hemolysis | |||
| Low toxicity |
| Matrix set | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|
| Learned | |||
| Random 1 | |||
| Random 2 | |||
| Random 3 |
| Matrix set | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|
| Learned | |||
| Random 1 | |||
| Random 2 | |||
| Random 3 |
| Matrix set | Activity score (%) | Non-hemolysis score (%) | Non-toxicity score (%) | Pooled HV |
|---|---|---|---|---|
| Learned | ||||
| Random 1 | ||||
| Random 2 | ||||
| Random 3 |
| Method | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|
| Search-policy baselines | |||
| Fixed coefficients | |||
| Random actions | |||
| LinUCB | |||
| NSGA-II | |||
| Proposed framework | |||
| Method | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|
| AMPGPT2 | |||
| AMPGAN | |||
| AMPGen | |||
| AMPDesigner | |||
| AMP-adapted DPLM | |||
| EASER |
| Method | Pareto sequences | Occupied cells | Coverage (%) |
|---|---|---|---|
| Search-policy baselines | |||
| Fixed coefficients | |||
| Random actions | |||
| LinUCB | |||
| NSGA-II | |||
| External generators | |||
| Proposal | Activity score | Non-hemolysis score | Non-toxicity score | Steer samples |
|---|---|---|---|---|
| 1 | ||||
| 2 | ||||
| 3 | ||||
| 4 |
| Score convention | Sequence | Activity score | Non-hemolysis score | Non-toxicity score |
|---|---|---|---|---|
| Before penalty | Parent | |||
| Before penalty | Offspring | |||
| After penalty | Parent | |||
| After penalty | Offspring |
| Control | Predicted CPP score | Score (%) |
|---|---|---|
| Zero | ||
| Matched random ( ) | ||
| Learned ( ) | 0.7067 ± 0.0220 | 76.8 ± 0.9 |
| Learned ( ) |