M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization
Authors: Junjie Wang, Yaowei Jin, Ruohui Tang, Guonan Cui, Haojie Wang, Penglei Wang, Dingyan Wang, Duo An, +2 more
Organizations: Lingang Laboratory, Shanghai, 200031, China. · School of Information Science and Technology, ShanghaiTech University, Shanghai, 201210, China. · Global Institute of Future Technology, Shanghai Jiao Tong University, Shanghai, 200240, China.
Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they must recover candidate identities, prior evaluations, and task constraints to guide subsequent decisions. We present M3OS, a multi-agent LLM system that decouples molecular-design reasoning from optimization-state management through Monte Carlo graph search. A persistent graph links evaluated candidates, parent-child transformations and evaluation evidence, while rewards and visit statistics guide LLM-assisted parent selection. Two branches combine tool-driven candidate generation with knowledge- and case-guided medicinal-chemistry editing. An execution harness controls graph updates through structured output extraction, molecular validation and task-bound evaluation. Agents receive role-specific contexts, while the graph preserves optimization trajectories beyond their active contexts. Across three molecular optimization benchmarks, M3OS achieves higher success rates than baselines, supporting the integration of persistent search state, specialized agents and controlled execution for multi-constraint optimization.
Figures & tables
Figure 1: M3OS architecture and optimization loop. a , Specialized agents share a persistent molecular graph that links candidates, structural transformations and evaluation evidence. b , Node selection, parallel Creative and Rational generation, molecular validation, common Critic evaluation and reward updates connect successive optimization rounds. Knowledge resources, tools and procedural skills support these operations.
QED
HBD
HBA
Sol.
BBBP
hERG
Sol.+HBA
QED+BBBP
Model
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
Direct general-purpose language models
GPT-5.5
0.94
0.79
0.92
0.28
0.85
0.32
0.83
0.37
0.86
0.67
0.66
0.32
0.65
0.17
0.68
0.37
Claude Opus 4.6
0.95
0.78
0.93
0.24
0.83
0.18
0.89
0.31
0.88
0.71
0.72
0.41
0.68
0.09
0.73
0.42
Gemini 3.5 Flash
0.87
0.69
0.93
0.14
0.88
0.25
0.86
0.38
0.87
0.66
0.61
0.29
0.74
0.18
0.71
0.42
Kimi-K3 Low
0.90
0.76
0.89
0.16
0.87
0.17
0.83
0.41
0.83
0.70
0.67
0.35
0.65
0.14
0.61
0.33
Table 1: MolOpt-120 success rates (0–1; 15 records per task). L/S: loose/strict success; HBA/HBD: hydrogen-bond acceptors/donors. Rates average candidates, then records; invalid candidates count as failures. M3OS uses up to 20 system-ranked, unique non-root candidates per cumulative graph; critic-score ties at the cutoff are resolved by graph insertion order. Baseline duplicates are retained. Solubility, BBBP and hERG use ADMET-AI 1.3.1. Bold/underline mark the best/second-best unrounded rates; ties share emphasis.
BPQ
MPQ
BHMQ
BMPQ
HMPQ
Model
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
Direct general-purpose language models
GPT-5.5
0.77
0.46
1.76
0.50
0.58
0.25
0.60
0.29
3.71
0.57
0.43
0.38
0.65
0.40
1.48
Claude Opus 4.6
0.73
0.51
1.68
0.65
0.58
0.45
0.62
0.35
4.24
0.58
0.49
0.80
0.70
0.42
1.93
Gemini 3.5 Flash
0.75
0.56
0.85
0.61
0.55
0.41
0.63
0.43
3.35
0.66
0.45
0.73
0.65
0.46
2.23
Kimi-K3 Low
0.50
0.57
0.51
0.61
0.54
0.36
0.59
0.43
3.22
0.51
0.46
0.66
0.66
0.37
1.87
Table 2: MuMOInstruct-100 results (20 records per task). Metrics average valid candidates, then records with valid outputs. SR: joint success rate (0–1); Sim.: Tanimoto similarity; RI: signed relative improvement. Task letters B/H/M/P/Q denote BBBP/HIA/mutagenicity/penalized logP/QED. M3OS uses up to 20 unique non-root candidates per cumulative graph, ranked in descending order by the system score; ties at the cutoff are resolved by graph insertion order. Baseline duplicates are retained. Bold/underline mark the best/second-best displayed values; ties share emphasis. Dashes indicate no valid candidates.
Model
Single solution SR ↑
Graph-any SR ↑
Kimi-K3 Low (Claude Code)
0.42
0.47
Kimi-K3 Low (Codex)
0.48
0.57
Kimi-K3 Low (M3OS R3)
0.59
0.61
Table 3: SMDD-Bench-93 task success rates. Single solution evaluates the final molecule; graph-any tests whether any recorded candidate satisfies all task constraints and Nesso activity criteria. SR is on a 0–1 scale. Bold/underline mark the best/second-best values.
Figure 2: Ablation study. a , MuMOInstruct SR, similarity and RI for Rational-only, Creative-only and complete M3OS at low reasoning effort, plus complete high-effort M3OS. SR: success rate; RI: relative improvement. b , Loose/strict success on BBBP/hERG (15 records each), all using Rational-only generation; “All” retains both retrieval sources. c , Ten-round complete low-effort M3OS: RI on the left axis, SR (%) on the right. Panels use cumulative oracle top-20 selection and equal task averaging; MuMOInstruct conditions on valid candidates.
Figure 3: Constrained molecular optimization on hDHFR. (a) Search graph and candidate counts for three optimization rounds. Purple circles denote task-qualified candidates; crossed red circles denote critic-rejected candidates. Highlighted edges trace the molecular-edit paths shown in (b) . (b) Representative molecular-edit trajectories, with modified regions highlighted and search-time Nesso scores and Tanimoto similarities shown below each structure. (c) Fraction of qualified complexes contacting each residue; labels give counts. (d) Predicted pocket views of mol15 and mol21. Ligands are yellow and orange, respectively, and selected pocket residues are cyan. Annotated heavy-atom distances connect the terminal hydroxyl oxygen to Arg70 Nη2 in mol15 and to the Phe31 backbone carbonyl oxygen in mol21. (e) Terminal hydroxyl–Arg70 distance grouped by distal-arm length and N -substitution series. Points represent molecules, vertical ticks indicate row medians, and right-hand labels give sample sizes. Residues use mature hDHFR numbering.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Primary Tool Components
Molecular and protein data retrieval
UniProt Consortium (2015) , ChEMBL Gaulton et al. (2012) , STOUT Rajan et al. (2024)
DuckDuckGo, PubMed, Semantic Scholar, MinerU Wang et al. (2024)
SMILES processing, validation, and topological graph generation
RDKit Bento et al. (2020)
Molecular generation
REINVENT4 Loeffler et al. (2024) with Mol2Mol Tibo et al. (2024) , LinkInvent Guo et al. (2023) , and LibInvent Fialková et al. (2021)
ADMET prediction and constraint screening
ADMET-AI Swanson et al. (2024) ; RDKit
Appendix
Table 4: Computational capabilities used by the current M3OS workflow.
Skill
Assigned agent(s)
Description
Candidate-quality gating
Critic
Evaluates all current-round candidates with frozen ADMET endpoints or explicit task constraints, then combines quantitative evidence with medicinal-chemistry risk review.
Nesso activity evaluation
Critic
Evaluates both generator branches with Nesso, adds comparison baselines, and preserves deterministic activity gates, scores and ranks.
Case-guided medicinal design
Rational
Converts pharmacophore and SAR hypotheses, medicinal-chemistry knowledge and retrieved optimization cases into interpretable, goal-calibrated molecular edits.
Medicinal-chemistry knowledge retrieval
Retrieval
Routes focused questions through shared memory, curated sources, the concept-oriented knowledge base and, when required, external literature.
Molecular optimization case retrieval
Rational
Retrieves scaffold-similar subgraphs from property-specific evolution graphs and extracts transferable transformation signals with their applicability limits.
Concept-oriented knowledge retrieval
Retrieval
Selects concept, article or multi-article retrieval according to query scope.
Appendix
Table 5: Active procedural skills and their role assignments in M3OS.
Function
Description
Local M3OS orchestration
get_uploaded_file_info
Documents and metadata uploaded to the session.
ask_medchem_knowledge
Concise, source-aware guidance from the MedChem Retrieval Agent.
run_mcgs_search
Graph search with the requested objective, generator mode, expansion-round limit and candidate target.
generate_optimization_report
Validated self-contained HTML decision report from the current graph and optimization trace.
screen_reinvent_candidates_by_constraints
Hard/ADMET screening of the current generation CSV; returns at most five passing Creative candidates.
Appendix
Table 6: M3OS orchestration functions and supporting MCP domain interfaces.
Figure 4: Knowledge and case resources. a , Source documents form an optimization-oriented collection and concept wiki. b , Matched molecular pairs form annotated optimization graphs. c , Scaffold matching retrieves connected transformation examples for rational design. Case-graph construction is detailed in Appendix 7 .
Property
Nodes
Edges
Roots
Components
Largest component
BBBP
674
665
331
155
35
hERG
1,934
2,651
1,063
360
54
Human liver microsomal stability
2,391
3,956
1,228
388
80
Apparent permeability ( Papp )
2,171
5,349
1,019
293
163
Appendix
Table 7: Property-specific molecular evolution graphs used by M3OS. Components are weakly connected components of the directed graph.
Figure 5: Search-report generation. Molecular depictions, predictions, rationales and an interactive graph are assembled from a graph snapshot into an HTML report. Validation gates delivery; PASS/FAIL indicate check outcomes.
Figure 6: Example of an M3OS optimization decision report. Selected views of the self-contained HTML report generated from the three-round search graph for the DHFR_16 task. a , Decision summary and reference-molecule profile. b , Interactive molecular search graph with node-level details and round-wise statistics. c , Focused candidate cards showing direct-parent structural comparisons with highlighted edits, predicted activity and ADMET properties, task-constraint checks, and generator and Critic rationales. d , Excerpt from the full-node audit, linking molecule identifiers to structures, complete canonical SMILES and recorded parent–child relationships. e , Recommendation tiers and a five-compound synthesis shortlist. f , Risk summaries and proposed follow-up experiments. Activity and ADMET values are computational predictions, whereas RDKit descriptors are deterministic calculations; recommendations remain subject to experimental validation.
Benchmark
Harness
Input
Output
MolOpt-120
Claude Code
61.03
1.33
MolOpt-120
Codex
123.91
1.12
MolOpt-120
M3OS
63.46
2.71
MuMOInstruct-100
Claude Code
57.20
1.31
MuMOInstruct-100
Codex
169.43
1.17
MuMOInstruct-100
M3OS
60.27
2.40
Appendix
Table 8: Aggregate token consumption across the three benchmarks. All values are in millions of tokens.
Molecular optimization, the process of designing molecules with desirable properties, represents a critical challenge in drug discovery. The recent advancements in large language models (LLMs) have opened new opportunities for their integration with traditional molecular optimization algorithms to improve performance. In this work, we propose Molecular Language Model powered Evolutionary Algorithm (Mol-E), an evolutionary algorithm that relies on the generative capabilities of LLMs trained on molecules and molecular properties. Scientific Contribution. Mol-E obtains the highest aggregate Top-10 AUC among the comparable full-23-task results considered here, scoring 17.500 in the task-agnostic regime, in which the oracle is treated strictly as a black box, and 20.551 in the task-informed regime, in which the optimizer receives a fixed semantic description of the objective. Mol-E also improves over the evaluated baselines on multi-property optimization with docking against DRD2, MK2, and AChE.
Philipp Guevorguian, Menua Bedrosian, Tigran Fahradyan +3
Large language models (LLMs) show promise for molecular optimization, but aligning them with selective and competing drug-design constraints remains challenging. We propose C-Moral, a reinforcement learning post-training framework for controllable multi-objective molecular optimization. C-Moral combines group-based relative optimization, property score alignment for heterogeneous objectives, and bottleneck-sensitive non-linear reward aggregation to improve stability across competing molecular properties. Experiments on C-MuMOInstruct and S2-Bench MolOpt show that C-Moral achieves the best performance among compared methods on both benchmarks. On C-MuMOInstruct, C-Moral achieves the best Success Optimized Rate (SOR) of 48.9% on in-domain tasks and 39.5% on out-of-domain tasks while preserving scaffold similarity. On S2-Bench MolOpt, it also achieves the strongest results across LogP, MR, and QED optimization tasks. These results suggest that C-Moral is an effective way to align molecular LLMs with continuous and constrained molecular design objectives. Our code and models are publicly available at https://github.com/Rwigie/C-MORAL.
Molecular optimization must improve target activity and satisfy developability constraints within limited evaluation budgets. Classical methods require tailored rules or training to incorporate chemical instructions and property feedback. Frozen language models can condition edits on this information, but need explicit constraint control and relevant experience. We therefore present PharmAgent, a constraint-aware molecular search method driven by adaptive external state. Its Lagrangian controller translates violations in accepted states into accumulated constraint pressure, keeping this history separate from current property measurements. Structure-indexed replay complements this feedback with relevant evaluated transitions that guide subsequent proposals. As a curriculum progressively activates constraints, candidates and the incumbent are compared under the same current objective, and the accepted state determines the next multiplier update. We derive an exact identity that characterizes how accepted-state violations accumulate in the controller's multipliers. Across five tasks with five independent runs, PharmAgent achieves a summed area under the target-score curves (AUC) of 3.9208 in target-only search, improving over MOLLEO by 37.3%. With online constraints, it achieves a property-adjusted AUC of 0.7076, improving over the strongest online baseline, ExLLM, by 53.8%. These results rank first among all evaluated methods in both target-only and constraint-aware search. The online comparison covers all five baseline frameworks. The full system leads every ablation variant in target quality, property-adjusted performance, and Pareto hypervolume. All five molecular cases reach feasible final states, documenting target gains and trade-offs.
Nihui Shao, Guanxing Chen, Jilong Shi +6
City University of Hong Kong (Dongguan) · The Hong Kong Polytechnic University · Tencent AI Lab +2