M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization
Authors: Junjie Wang, Yaowei Jin, Ruohui Tang, Guonan Cui, Haojie Wang, Penglei Wang, Dingyan Wang, Duo An, +2 more
Organizations: Lingang Laboratory, Shanghai, 200031, China. · School of Information Science and Technology, ShanghaiTech University, Shanghai, 201210, China. · Global Institute of Future Technology, Shanghai Jiao Tong University, Shanghai, 200240, China.
Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they must recover candidate identities, prior evaluations, and task constraints to guide subsequent decisions. We present M3OS, a multi-agent LLM system that decouples molecular-design reasoning from optimization-state management through Monte Carlo graph search. A persistent graph links evaluated candidates, parent-child transformations and evaluation evidence, while rewards and visit statistics guide LLM-assisted parent selection. Two branches combine tool-driven candidate generation with knowledge- and case-guided medicinal-chemistry editing. An execution harness controls graph updates through structured output extraction, molecular validation and task-bound evaluation. Agents receive role-specific contexts, while the graph preserves optimization trajectories beyond their active contexts. Across three molecular optimization benchmarks, M3OS achieves higher success rates than baselines, supporting the integration of persistent search state, specialized agents and controlled execution for multi-constraint optimization.
Figures & tables
Figure 1: M3OS architecture and optimization loop. a , Specialized agents share a persistent molecular graph that links candidates, structural transformations and evaluation evidence. b , Node selection, parallel Creative and Rational generation, molecular validation, common Critic evaluation and reward updates connect successive optimization rounds. Knowledge resources, tools and procedural skills support these operations.
QED
HBD
HBA
Sol.
BBBP
hERG
Sol.+HBA
QED+BBBP
Model
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
L ↑
S ↑
Direct general-purpose language models
GPT-5.5
0.94
0.79
0.92
0.28
0.85
0.32
0.83
0.37
0.86
0.67
0.66
0.32
0.65
0.17
0.68
0.37
Claude Opus 4.6
0.95
0.78
0.93
0.24
0.83
0.18
0.89
0.31
0.88
0.71
0.72
0.41
0.68
0.09
0.73
0.42
Gemini 3.5 Flash
0.87
0.69
0.93
0.14
0.88
0.25
0.86
0.38
0.87
0.66
0.61
0.29
0.74
0.18
0.71
0.42
Kimi-K3 Low
0.90
0.76
0.89
0.16
0.87
0.17
0.83
0.41
0.83
0.70
0.67
0.35
0.65
0.14
0.61
0.33
Table 1: MolOpt-120 success rates (0–1; 15 records per task). L/S: loose/strict success; HBA/HBD: hydrogen-bond acceptors/donors. Rates average candidates, then records; invalid candidates count as failures. M3OS uses up to 20 system-ranked, unique non-root candidates per cumulative graph; critic-score ties at the cutoff are resolved by graph insertion order. Baseline duplicates are retained. Solubility, BBBP and hERG use ADMET-AI 1.3.1. Bold/underline mark the best/second-best unrounded rates; ties share emphasis.
BPQ
MPQ
BHMQ
BMPQ
HMPQ
Model
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
SR ↑
Sim. ↑
RI ↑
Direct general-purpose language models
GPT-5.5
0.77
0.46
1.76
0.50
0.58
0.25
0.60
0.29
3.71
0.57
0.43
0.38
0.65
0.40
1.48
Claude Opus 4.6
0.73
0.51
1.68
0.65
0.58
0.45
0.62
0.35
4.24
0.58
0.49
0.80
0.70
0.42
1.93
Gemini 3.5 Flash
0.75
0.56
0.85
0.61
0.55
0.41
0.63
0.43
3.35
0.66
0.45
0.73
0.65
0.46
2.23
Kimi-K3 Low
0.50
0.57
0.51
0.61
0.54
0.36
0.59
0.43
3.22
0.51
0.46
0.66
0.66
0.37
1.87
Table 2: MuMOInstruct-100 results (20 records per task). Metrics average valid candidates, then records with valid outputs. SR: joint success rate (0–1); Sim.: Tanimoto similarity; RI: signed relative improvement. Task letters B/H/M/P/Q denote BBBP/HIA/mutagenicity/penalized logP/QED. M3OS uses up to 20 unique non-root candidates per cumulative graph, ranked in descending order by the system score; ties at the cutoff are resolved by graph insertion order. Baseline duplicates are retained. Bold/underline mark the best/second-best displayed values; ties share emphasis. Dashes indicate no valid candidates.
Model
Single solution SR ↑
Graph-any SR ↑
Kimi-K3 Low (Claude Code)
0.42
0.47
Kimi-K3 Low (Codex)
0.48
0.57
Kimi-K3 Low (M3OS R3)
0.59
0.61
Table 3: SMDD-Bench-93 task success rates. Single solution evaluates the final molecule; graph-any tests whether any recorded candidate satisfies all task constraints and Nesso activity criteria. SR is on a 0–1 scale. Bold/underline mark the best/second-best values.
Figure 2: Ablation study. a , MuMOInstruct SR, similarity and RI for Rational-only, Creative-only and complete M3OS at low reasoning effort, plus complete high-effort M3OS. SR: success rate; RI: relative improvement. b , Loose/strict success on BBBP/hERG (15 records each), all using Rational-only generation; “All” retains both retrieval sources. c , Ten-round complete low-effort M3OS: RI on the left axis, SR (%) on the right. Panels use cumulative oracle top-20 selection and equal task averaging; MuMOInstruct conditions on valid candidates.
Figure 3: Constrained molecular optimization on hDHFR. (a) Search graph and candidate counts for three optimization rounds. Purple circles denote task-qualified candidates; crossed red circles denote critic-rejected candidates. Highlighted edges trace the molecular-edit paths shown in (b) . (b) Representative molecular-edit trajectories, with modified regions highlighted and search-time Nesso scores and Tanimoto similarities shown below each structure. (c) Fraction of qualified complexes contacting each residue; labels give counts. (d) Predicted pocket views of mol15 and mol21. Ligands are yellow and orange, respectively, and selected pocket residues are cyan. Annotated heavy-atom distances connect the terminal hydroxyl oxygen to Arg70 Nη2 in mol15 and to the Phe31 backbone carbonyl oxygen in mol21. (e) Terminal hydroxyl–Arg70 distance grouped by distal-arm length and N -substitution series. Points represent molecules, vertical ticks indicate row medians, and right-hand labels give sample sizes. Residues use mature hDHFR numbering.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Primary Tool Components
Molecular and protein data retrieval
UniProt Consortium (2015) , ChEMBL Gaulton et al. (2012) , STOUT Rajan et al. (2024)
DuckDuckGo, PubMed, Semantic Scholar, MinerU Wang et al. (2024)
SMILES processing, validation, and topological graph generation
RDKit Bento et al. (2020)
Molecular generation
REINVENT4 Loeffler et al. (2024) with Mol2Mol Tibo et al. (2024) , LinkInvent Guo et al. (2023) , and LibInvent Fialková et al. (2021)
ADMET prediction and constraint screening
ADMET-AI Swanson et al. (2024) ; RDKit
Appendix
Table 4: Computational capabilities used by the current M3OS workflow.
Skill
Assigned agent(s)
Description
Candidate-quality gating
Critic
Evaluates all current-round candidates with frozen ADMET endpoints or explicit task constraints, then combines quantitative evidence with medicinal-chemistry risk review.
Nesso activity evaluation
Critic
Evaluates both generator branches with Nesso, adds comparison baselines, and preserves deterministic activity gates, scores and ranks.
Case-guided medicinal design
Rational
Converts pharmacophore and SAR hypotheses, medicinal-chemistry knowledge and retrieved optimization cases into interpretable, goal-calibrated molecular edits.
Medicinal-chemistry knowledge retrieval
Retrieval
Routes focused questions through shared memory, curated sources, the concept-oriented knowledge base and, when required, external literature.
Molecular optimization case retrieval
Rational
Retrieves scaffold-similar subgraphs from property-specific evolution graphs and extracts transferable transformation signals with their applicability limits.
Concept-oriented knowledge retrieval
Retrieval
Selects concept, article or multi-article retrieval according to query scope.
Appendix
Table 5: Active procedural skills and their role assignments in M3OS.
Function
Description
Local M3OS orchestration
get_uploaded_file_info
Documents and metadata uploaded to the session.
ask_medchem_knowledge
Concise, source-aware guidance from the MedChem Retrieval Agent.
run_mcgs_search
Graph search with the requested objective, generator mode, expansion-round limit and candidate target.
generate_optimization_report
Validated self-contained HTML decision report from the current graph and optimization trace.
screen_reinvent_candidates_by_constraints
Hard/ADMET screening of the current generation CSV; returns at most five passing Creative candidates.
Appendix
Table 6: M3OS orchestration functions and supporting MCP domain interfaces.
Figure 4: Knowledge and case resources. a , Source documents form an optimization-oriented collection and concept wiki. b , Matched molecular pairs form annotated optimization graphs. c , Scaffold matching retrieves connected transformation examples for rational design. Case-graph construction is detailed in Appendix 7 .
Property
Nodes
Edges
Roots
Components
Largest component
BBBP
674
665
331
155
35
hERG
1,934
2,651
1,063
360
54
Human liver microsomal stability
2,391
3,956
1,228
388
80
Apparent permeability ( Papp )
2,171
5,349
1,019
293
163
Appendix
Table 7: Property-specific molecular evolution graphs used by M3OS. Components are weakly connected components of the directed graph.
Figure 5: Search-report generation. Molecular depictions, predictions, rationales and an interactive graph are assembled from a graph snapshot into an HTML report. Validation gates delivery; PASS/FAIL indicate check outcomes.
Figure 6: Example of an M3OS optimization decision report. Selected views of the self-contained HTML report generated from the three-round search graph for the DHFR_16 task. a , Decision summary and reference-molecule profile. b , Interactive molecular search graph with node-level details and round-wise statistics. c , Focused candidate cards showing direct-parent structural comparisons with highlighted edits, predicted activity and ADMET properties, task-constraint checks, and generator and Critic rationales. d , Excerpt from the full-node audit, linking molecule identifiers to structures, complete canonical SMILES and recorded parent–child relationships. e , Recommendation tiers and a five-compound synthesis shortlist. f , Risk summaries and proposed follow-up experiments. Activity and ADMET values are computational predictions, whereas RDKit descriptors are deterministic calculations; recommendations remain subject to experimental validation.
Benchmark
Harness
Input
Output
MolOpt-120
Claude Code
61.03
1.33
MolOpt-120
Codex
123.91
1.12
MolOpt-120
M3OS
63.46
2.71
MuMOInstruct-100
Claude Code
57.20
1.31
MuMOInstruct-100
Codex
169.43
1.17
MuMOInstruct-100
M3OS
60.27
2.40
Appendix
Table 8: Aggregate token consumption across the three benchmarks. All values are in millions of tokens.