A self-learning scientific agent for X-ray diffraction
Authors: Bin Cao, Huichi Zhou, Runyu Yang, Jingsong Li, Shuchen Sun, Yan Song, Hanyu Gao, Zhongwei Yu, +2 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou), China · University College London, United Kingdom · AI Lab, The Yangtze River Delta, China · Institute of Automation, Chinese Academy of Sciences, China
A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30%, 81.78% and 40.83% on MP500, RRUFF and opXRD, respectively, compared with 58.00%, 58.47% and 26.45% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.
Figures & tables
Figure 1: Diverse materials applications and leading diffraction-identification performance. a , Whole-pattern refinement of orthorhombic PbSOX4 with strongly overlapping reflections ( Rp=2.753% , Rwp=5.926% ). b , Five-phase analysis of an ancient Egyptian cosmetic powder containing gypsum, phosgenite, cerussite, galena and laurionite ( Rp=6.600% , Rwp=13.079% ). c , Operando analysis of an O3-type Li(NiX0.8CoX0.1MnX0.1)OX2 cathode. The voltage–time trace (left), diffraction waterfall (centre; colour indicates scan order) and refined c axis (right) connect electrochemical cycling to lattice evolution. d , Refinement of the selected Ru–Mn oxide configuration ( Rp=1.447% , Rwp=3.129% ). In a,b,d , pale blue symbols and red curves denote observed and calculated intensities; ΔI=Iobs−Icalc is the difference profile, and ticks mark calculated Bragg positions. e , Single-phase identification on MP500, RRUFF and opXRD: bars show top-1 accuracy and heat maps show top-1, top-3, top-5 and mean reciprocal rank (MRR@5). f , Multiphase identification on the same sources: bars show sample-wise macro-F1 and heat maps show prediction coverage (Cov.), exact-set recovery (Exact), precision (P), recall (R) and F1. Open and filled bars denote XRD-only and composition-assisted conditions. Gan Jiang is highlighted in red. Higher scores are better. Full values and composition-input protocols are given in Tables 1 and 2 .
Figure 2: Diffraction physics, analytical workflow and reusable learning in Gan Jiang. a , A crystal’s periodic atomic arrangement produces diffraction at angles determined by lattice spacings. Randomly oriented crystallites generate a powder pattern whose peak positions, relative intensities and widths provide information about the lattice, atomic arrangement and microstructure, subject to instrumental effects. b , Gan Jiang links five analysis stages: input, candidate retrieval, evidence checking, refinement and reporting. The measured pattern is combined with available chemical and experimental context. XMatcher and XQueryer-LW propose single-phase candidates; XDecomposer and XMatcher-AutoMix support mixture analysis when single-phase hypotheses fail the evidence checks. Candidate structures are assessed against peak agreement, chemistry and coverage, then refined with eligible engines. Outputs include CIFs, observed and calculated profiles, refined parameters and an evidence-linked report. The central loop connects scientific constraints, analytical decisions and learning from completed episodes. c , Distribution of atoms per unit cell in the RRUFF and combined opXRD/private branches of the experimental structure collection, expressed as percentages of records within each branch. d , Representative PbSOX4 whole-pattern fit. e , Summary of FullProf workflow revision, illustrating the accumulation of improvements. Development trajectories and frozen-version held-out comparisons are distinguished in Fig. 4 .
Figure 3: Executable architecture for evidence-governed diffraction analysis. The user objective and constraints determine the skill and analysis plan; state, policy and capability checks govern execution. a , Local parsing and validation produce an observed pattern (O; full range and original scale), a model input (M; 3,500 normalized points over 10∘ – 90∘ ) and an inspector input (I; full range in the required format). b , XMatcher retrieves single-phase candidates; user-supplied elements enable XQueryer-LW. Candidates are merged while retaining identities and source ranks. XMatcher-AutoMix and XDecomposer support requested mixture analysis or automatic fallback after rejection of single-phase hypotheses. c , Candidate CIFs are exported and inspected against the measurement. Finalization requires the inspector verdict, elemental consistency where applicable and mandatory peak-evidence checks. Eligible angular-shift trials supplement the unshifted comparison. d , Supported phase sets or user-provided CIFs enter refinement through PyXplore (the PyWPEM service), FullProf or GSAS-II. An inspector checks the resulting profiles; passing runs are selected by inspection score and then peak coverage. At most one eligible budget-related retry is allowed. Solid arrows indicate data or evidence flow; dashed arrows indicate conditional control, including rejection and retry paths.
Figure 4: Self-improvement of refinement workflows and transfer to held-out samples. a–c , Learning trajectories for FullProf workflow revision, GSAS-II joint optimization and PyWPEM skill optimization (labelled WPEM), respectively. Upper plots show cumulative development-score gains; lower step plots show the increments associated with revisions. Highlighted intervals identify the large increments marked in each panel. Comparisons beneath the trajectories report mean held-out scores for the frozen initial and final workflows: 46.90 to 72.38 for FullProf, 39.01 to 60.02 for GSAS-II and 31.74 to 38.33 for PyWPEM. d , Controlled learning protocol. Training episodes provide execution and diffraction feedback for revising instructions, code and parameters. Paired development checks select candidate revisions. The selected version is frozen before held-out evaluation, and test results do not select the learning trajectory. e , Deployment-time learning architecture. New user episodes are analysed with the current skill library; verified tool evidence and user feedback inform candidate revisions. Regression checks precede versioned deployment, with monitoring and rollback.
XRD-only
XRD + composition
Model
Top-1 ↑ (%)
Top-3 ↑ (%)
Top-5 ↑ (%)
MRR@5 ↑ (%)
Top-1 ↑ (%)
Top-3 ↑ (%)
Top-5 ↑ (%)
MRR@5 ↑ (%)
MP500 ( N=10,000 )
Uni3DAR
0.02
0.02
0.02
0.02
29.24
30.89
31.16
30.07
PXRDGen
0.01
0.01
0.01
0.01
18.79
24.58
26.53
21.82
XtalNet
0.00
0.00
0.00
0.00
0.00
0.00
0.20
0.04
AutoXRD
58.00
61.90
63.60
60.17
96.20
97.90
98.30
97.04
Table 1: Single-phase structure identification across simulated and experimental diffraction patterns. Results are reported for end-to-end XRD-only inference and XRD + oracle composition. All metrics are higher-is-better; the best and second-best values within each dataset and condition are shown in bold and underlined , respectively. AutoXRD uses a fixed candidate library (GanJiang dataset). PXRDGen-Flow’s reduced RRUFF and opXRD subsets satisfy its ≤100 -atom limit. PXRDGen and XtalNet abbreviate PXRDGen-Flow and XtalNet-HMOF100 (944 RRUFF and 654 opXRD).
XRD-only
XRD + composition
Model
Cov. ↑ (%)
Exact ↑ (%)
P ↑ (%)
R ↑ (%)
F1 ↑ (%)
Cov. ↑ (%)
Exact ↑ (%)
P ↑ (%)
R ↑ (%)
F1 ↑ (%)
MP500 ( N=1,000 )
AutoAnalyzer
100.00
1.70
46.49
58.08
51.12
95.63
23.37
85.08
58.08
69.25
S-CNN
100.00
1.27
40.11
18.99
25.06
43.17
1.60
42.46
18.99
25.72
AutoXRD
92.40
15.80
50.22
40.48
43.24
98.50
59.00
85.57
80.12
81.75
Gan Jiang
100.00
22.80
76.85
64.77
66.98
100.00
69.70
94.97
89.68
91.61
Table 2: Multiphase identification across simulated and experimental diffraction patterns. Each dataset contains 1,000 hidden mixtures. Results are reported under XRD-only and composition-assisted conditions. Higher values are better for all metrics; the best and second-best values within each dataset and condition are shown in bold and underlined , respectively. Cov. denotes the percentage of mixtures with at least one valid predicted phase after parsing, deduplication and any declared composition filtering. Exact denotes complete phase-set recovery. Exact, precision (P), recall (R), and F1 are evaluated over all 1,000 mixtures in each run. P, R, and F1 are sample-wise macro averages.
Track
Required prediction
Primary metrics
Single-phase identification
Ranked list of at most five CIFs
Top-1, Top-3, Top-5; MRR@5
Multi-phase identification
Unordered CIF set
coverage; exact set; macro precision, recall and F1
Refinement
Calculated intensity pattern
Rp , Rwp , correlation, score
Table S1: DeltaXRDbench task contract. A phase set is evaluated after one-to-one matching between submitted and reference structures.
Figure S1: Representative single- and multiphase samples in DeltaXRDbench. a,b , Simulated MP500 patterns; c,d , experimental RRUFF patterns; and e,f , experimental opXRD patterns. a,c,e show representative single-phase examples, pairing each normalized powder-XRD profile with its crystal structure. b,d,f show representative multiphase mixtures, with the combined profile at left and the constituent phase structures at right. The horizontal axis is 2θ (degrees) and the vertical axis is normalized intensity; labels identify the corresponding database records. Sphere colours distinguish atomic elements and dashed polyhedra indicate unit-cell boundaries.
Source
Single
Multi
Total
MP500 (simulated)
10,000
30,000
40,000
RRUFF (experimental)
1,164
10,000
11,164
opXRD (experimental)
880
10,000
10,880
All sources
12,044
50,000
62,044
Table S2: Dataset composition in the current DeltaXRDbench release. “Single” denotes single-phase patterns, whereas “Multi” denotes multiphase patterns generated by mixing single-phase patterns.
Quantity
Range / setting
Grain size
10–100
Preferred orientation (each component)
−0.3 – 0.3
Thermal vibration
0–2
Zero shift
0–2
Detector–sample distance
300–600
Detector half-height / sample half-height
3–8 / 1–4
Table S3: Default ranges sampled for dynamic Pysimxrd simulation.
Component
Setting
Input grid
3,500 samples; 10–90 ∘
FFT views
original + 3 low-pass views
CNN base channels
64
Attention dimension / heads
192 / 6
XRD tokens / element queries
96 / 4
Attention refinement blocks
2
Table S4: Default model configuration. The parameter count was computed from count_parameters(Xmodel()) .
Figure S2: XDecomposer learns to separate multiphase diffraction signals without a predefined phase list. a , Two-stage learning. Stage I pretrains a global encoder by reconstructing masked regions of single-phase patterns. Stage II combines a hierarchical encoder, the frozen global encoder (snowflake), a feature-wise linear modulation (FiLM) phase guide and a symmetric decoder to produce a set of candidate component patterns from a mixed XRD input. b , Inference framework. Hierarchical multi-scale features and global context are conditioned by learned phase queries through cross-attention and aggregation; the decoder outputs sigmoid masks mk , which are applied to the input profile x to yield physics-consistent component candidates y^k . Curves represent diffraction intensity as a function of 2θ ; colours distinguish predicted components.
Category
Parameter
Value or range
Random perturbations
Crystallite size (nm)
10–120
Thermal vibration coefficient
0.01–0.20
Zero shift
0–0.20
Detector–sample distance (mm)
300–600
Detector slit half-height (mm)
3–8
Sample half-height (mm)
1–4
Table S5: Parameters used to generate simulated pretraining data for XDecomposer.
Figure S3: XRDinspector provides transparent quality assessment for PXRD results. a , Candidate CIF(s) and an experimental pattern are aligned to quantify peak coverage, precision, phase support and intensity agreement. b , Observed and calculated profiles are evaluated using Rp , Rwp , correlation, overlap and local-residual diagnostics. c , The unified policy converts these diagnostic scores into four delivery levels, from high-confidence delivery (L1) to rejection (L4), and returns an auditable report with metrics, warnings and settings.
Component
Role
Experiment
Provides diffraction data, structural candidates and instrument metadata.
Gan Jiang
Audits inputs, defines hypotheses, evaluates outputs and reports conclusions.
Automation
Isolates each run, executes the backend and records logs.
FullProf
Calculates diffraction profiles and performs least-squares refinement.
Table S6: Roles in the Gan Jiang–FullProf refinement workflow.
Figure S4: Evidence-governed refinement workflow for the FullProf and GSAS-II branches. Experimental XRD data, structural CIFs or PCR files, and instrument records first undergo language-model preflight checks of links, units, phase model and radiation assumptions. Insufficient evidence triggers a request for calibration or phase-identification information. The FullProf branch (top) copies inputs into a new PCR file, performs refinement and either preserves the resulting profiles, logs and JSON summary or proceeds to the next minimal parameter-group test. The GSAS-II branch (bottom) validates its staged inputs, executes Rietveld refinement, and subjects the result to physical and statistical review before recording project files, metrics and tracebacks. Yellow chevrons denote the ordered refinement stages; diamonds denote decision gates. Accepted analyses retain diagnostics, limitations and a qualified conclusion.
Table S7: Programmatic access to selected capabilities of the Gan Jiang diffraction toolchain. Endpoint identifiers are supplied to the --endpoint argument of delta-cli science invoke --tool xrd .
Figure S5: Compositional and structural landscape of the MP500 dataset. a , Principal-component map of elemental composition for all 100,315 stored structures; colour denotes the number of structures per hexagonal bin. b , Co-occurrence network of the 16 most prevalent elements; links represent the 28 largest pairwise co-occurrence counts. c , Distribution of chemical complexity, expressed as the number of distinct elements per structure. d , Empirical cumulative distribution of atoms per stored unit cell; vertical lines mark the median and 90th percentile. e , Pairwise co-occurrence matrix of prevalent elements.
Figure S6: Overview of the experimental candidate library. a , Distribution of the number of atoms in the stored unit cells; solid outlines show binned fractions, dashed curves show empirical cumulative distributions, and dotted vertical lines mark the medians. b , Crystal-system distributions within RRUFF and the combined opXRD + private-data branch; symmetries were identified with spglib using symprec=10−2A˚ , and numbers above the bars give record counts. c , Occurrence of the 12 most prevalent elements, with each element counted at most once per structure record. d , Source composition of the collection, with distinct CIF counts determined from the stored SHA-256 checksums.
Figure S7: Five-dimensional index of the MP500 XRD archive. a , Archive-wide occupancy of 2,934,906 stored theoretical reflection positions, with a modal value at 30.9 ∘ 2 θ . b , Structure-resolved diffraction barcode for 50 fixed-seed records; colour denotes normalized reflection intensity. c , Rank-frequency distribution of leading space groups. d , Hexagonally binned relation between stored-peak count and Shannon entropy of normalized peak intensity. e , Elemental support of the inverted index; colour denotes the number of structures containing an element.
Phase
a (Å)
b (Å)
c (Å)
α ( ∘ )
β ( ∘ )
γ ( ∘ )
Gypsum
5.6800(9)
15.2144(0)
6.5303(7)
90
118.4841(5)
90
Phosgenite
8.1600(0)
8.1600(0)
8.8834(3)
90
90
90
Cerussite
5.1794(7)
8.4922(9)
6.1418(1)
90
90
90
Galena
5.9388(0)
5.9388(0)
5.9388(0)
90
90
90
Laurionite
9.7002(6)
4.0200(3)
7.1108(1)
90
90
90
Table S8: Refined lattice parameters of the five mineral phases identified in the ancient Egyptian cosmetic powder. Values in parentheses are the estimated uncertainties in the final reported digit.
Figure S8: Whole-pattern decomposition of binary and ternary powder mixtures. a , Fits to a nominal 2:8 SiOX2 ( α -quartz)– TiOX2 (rutile) mixture (top) and a nominal 1:1:1 AlX2OX3 (corundum)– ZnO (wurtzite)– NaCl (rock salt) mixture (bottom). Black curves show the measured scans, grey dashed curves show the fitted background and coloured fills show the phase-resolved contributions. b , Component diffraction profiles recovered for each phase in the two mixtures, using the same colour code.
As scientific workflows shift from deterministic executables to LLM-based agents, the development practices on offer, such as fine-tuning, reinforcement learning, and prompt-and-go, bury the scientist's judgment. We propose treating agent construction as a workflow stage and introduce AgentBuild, which builds a scientific agent from a contract the scientist authors. The contract is a version-controlled rubric, a difficulty-graded curriculum, and a curated external knowledge base. A rubric-driven judge gates a meta-optimizer coding agent that edits the agent within a declared boundary, so the build compiles the agent, not the scientist's judgment. We instantiate this for Rietveld refinement of X-ray diffraction data through GSAS-II behind MCP and A2A, where a blank-harness construction run progresses through a lithium lanthanum zirconium oxide (LLZO) signal-to-noise ladder, reaches the 4 hour scan as a frontier case, and exposes the workflow-scope limits that remain. The same rubric that rewards credible fits also scores trajectory scope, making the frontier a contract failure rather than a pattern-fitting failure. As base models evolve, re-running AgentBuild is a re-tune, not a rebuild, and the scientist's authored contract remains the durable asset.
Woong Shin, Craig A. Bridges, Marshall T. McDonnell +1
Multiphase powder X-ray diffraction (PXRD) analysis remains a fundamental bottleneck in structure identification, as real-world synthesis often produces complex mixtures whose constituent phases (components) cannot be reliably disentangled. While recent advances in representation-based crystal retrieval and generation suggest the possibility of inferring structures directly from PXRD, existing approaches largely assume single-phase inputs and break down in multiphase settings. Here, we present XDecomposer, a prior-free framework for joint decomposition and identification of multiphase XRD patterns without requiring candidate phase lists, structural templates, or prior knowledge of phase number. We formulate multiphase diffraction analysis as a set prediction problem, where the model infers an unordered set of phase-resolved components, their mixture proportions, and corresponding structural representations within a unified architecture. A phase-query-driven decomposition mechanism, together with diffraction-consistent physical reconstruction, enables accurate source separation while preserving crystallographic fidelity. Extensive experiments on both simulated and experimental datasets show that XDecomposer substantially improves reconstruction accuracy and phase identification across diverse chemical systems, while maintaining strong generalization to unseen mixtures. These results provide a practical route toward data-driven, source-resolved multiphase XRD analysis and reduce long-standing dependence on prior-guided iteratively phase matching. The code is openly available at https://github.com/Licht0812/XDecomposer
Hanyu Gao, Bin Cao, Yunyue Su +2
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA) · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Guangzhou Municipal Key Laboratory of Materials Informatics, The Hong Kong University of Science and Technology (Guangzhou) +1
A language-model agent asked to analyse an experiment will usually return working code. Whether the analysis is defensible is a different question. A defensible analysis depends on procedural choices: which test the field accepts, which identifier namespace is authoritative, and which caveats must accompany a result. We present Scientific Agent Skills, an open library of 163 such procedures in 16 areas of practice, including genomics, cheminformatics, medical imaging, study design and scientific communication. Each skill is a directory built around a versioned, human-readable instruction file. An agent loads the file only when a task calls for it; the directory often also contains reference material and runnable scripts. We report no task-level evaluation and no host selection rate. We measure two properties of the documentation corpus: the always-resident descriptions of all 163 skills cost 7.1% of a 200,000-token window, and the median documented workflow fits within 23.9% of it, although 29 of 46 would overflow if every reference file were loaded. Openly licensed and available at https://github.com/K-Dense-AI/scientific-agent-skills.