A self-learning scientific agent for X-ray diffraction
Authors: Bin Cao, Huichi Zhou, Runyu Yang, Jingsong Li, Shuchen Sun, Yan Song, Hanyu Gao, Zhongwei Yu, +2 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou), China · University College London, United Kingdom · AI Lab, The Yangtze River Delta, China · Institute of Automation, Chinese Academy of Sciences, China
A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30%, 81.78% and 40.83% on MP500, RRUFF and opXRD, respectively, compared with 58.00%, 58.47% and 26.45% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.
Figures & tables
Figure 1: Diverse materials applications and leading diffraction-identification performance. a , Whole-pattern refinement of orthorhombic PbSOX4 with strongly overlapping reflections ( Rp=2.753% , Rwp=5.926% ). b , Five-phase analysis of an ancient Egyptian cosmetic powder containing gypsum, phosgenite, cerussite, galena and laurionite ( Rp=6.600% , Rwp=13.079% ). c , Operando analysis of an O3-type Li(NiX0.8CoX0.1MnX0.1)OX2 cathode. The voltage–time trace (left), diffraction waterfall (centre; colour indicates scan order) and refined c axis (right) connect electrochemical cycling to lattice evolution. d , Refinement of the selected Ru–Mn oxide configuration ( Rp=1.447% , Rwp=3.129% ). In a,b,d , pale blue symbols and red curves denote observed and calculated intensities; ΔI=Iobs−Icalc is the difference profile, and ticks mark calculated Bragg positions. e , Single-phase identification on MP500, RRUFF and opXRD: bars show top-1 accuracy and heat maps show top-1, top-3, top-5 and mean reciprocal rank (MRR@5). f , Multiphase identification on the same sources: bars show sample-wise macro-F1 and heat maps show prediction coverage (Cov.), exact-set recovery (Exact), precision (P), recall (R) and F1. Open and filled bars denote XRD-only and composition-assisted conditions. Gan Jiang is highlighted in red. Higher scores are better. Full values and composition-input protocols are given in Tables 1 and 2 .
Figure 2: Diffraction physics, analytical workflow and reusable learning in Gan Jiang. a , A crystal’s periodic atomic arrangement produces diffraction at angles determined by lattice spacings. Randomly oriented crystallites generate a powder pattern whose peak positions, relative intensities and widths provide information about the lattice, atomic arrangement and microstructure, subject to instrumental effects. b , Gan Jiang links five analysis stages: input, candidate retrieval, evidence checking, refinement and reporting. The measured pattern is combined with available chemical and experimental context. XMatcher and XQueryer-LW propose single-phase candidates; XDecomposer and XMatcher-AutoMix support mixture analysis when single-phase hypotheses fail the evidence checks. Candidate structures are assessed against peak agreement, chemistry and coverage, then refined with eligible engines. Outputs include CIFs, observed and calculated profiles, refined parameters and an evidence-linked report. The central loop connects scientific constraints, analytical decisions and learning from completed episodes. c , Distribution of atoms per unit cell in the RRUFF and combined opXRD/private branches of the experimental structure collection, expressed as percentages of records within each branch. d , Representative PbSOX4 whole-pattern fit. e , Summary of FullProf workflow revision, illustrating the accumulation of improvements. Development trajectories and frozen-version held-out comparisons are distinguished in Fig. 4 .
Figure 3: Executable architecture for evidence-governed diffraction analysis. The user objective and constraints determine the skill and analysis plan; state, policy and capability checks govern execution. a , Local parsing and validation produce an observed pattern (O; full range and original scale), a model input (M; 3,500 normalized points over 10∘ – 90∘ ) and an inspector input (I; full range in the required format). b , XMatcher retrieves single-phase candidates; user-supplied elements enable XQueryer-LW. Candidates are merged while retaining identities and source ranks. XMatcher-AutoMix and XDecomposer support requested mixture analysis or automatic fallback after rejection of single-phase hypotheses. c , Candidate CIFs are exported and inspected against the measurement. Finalization requires the inspector verdict, elemental consistency where applicable and mandatory peak-evidence checks. Eligible angular-shift trials supplement the unshifted comparison. d , Supported phase sets or user-provided CIFs enter refinement through PyXplore (the PyWPEM service), FullProf or GSAS-II. An inspector checks the resulting profiles; passing runs are selected by inspection score and then peak coverage. At most one eligible budget-related retry is allowed. Solid arrows indicate data or evidence flow; dashed arrows indicate conditional control, including rejection and retry paths.
Figure 4: Self-improvement of refinement workflows and transfer to held-out samples. a–c , Learning trajectories for FullProf workflow revision, GSAS-II joint optimization and PyWPEM skill optimization (labelled WPEM), respectively. Upper plots show cumulative development-score gains; lower step plots show the increments associated with revisions. Highlighted intervals identify the large increments marked in each panel. Comparisons beneath the trajectories report mean held-out scores for the frozen initial and final workflows: 46.90 to 72.38 for FullProf, 39.01 to 60.02 for GSAS-II and 31.74 to 38.33 for PyWPEM. d , Controlled learning protocol. Training episodes provide execution and diffraction feedback for revising instructions, code and parameters. Paired development checks select candidate revisions. The selected version is frozen before held-out evaluation, and test results do not select the learning trajectory. e , Deployment-time learning architecture. New user episodes are analysed with the current skill library; verified tool evidence and user feedback inform candidate revisions. Regression checks precede versioned deployment, with monitoring and rollback.
XRD-only
XRD + composition
Model
Top-1 ↑ (%)
Top-3 ↑ (%)
Top-5 ↑ (%)
MRR@5 ↑ (%)
Top-1 ↑ (%)
Top-3 ↑ (%)
Top-5 ↑ (%)
MRR@5 ↑ (%)
MP500 ( N=10,000 )
Uni3DAR
0.02
0.02
0.02
0.02
29.24
30.89
31.16
30.07
PXRDGen
0.01
0.01
0.01
0.01
18.79
24.58
26.53
21.82
XtalNet
0.00
0.00
0.00
0.00
0.00
0.00
0.20
0.04
AutoXRD
58.00
61.90
63.60
60.17
96.20
97.90
98.30
97.04
Table 1: Single-phase structure identification across simulated and experimental diffraction patterns. Results are reported for end-to-end XRD-only inference and XRD + oracle composition. All metrics are higher-is-better; the best and second-best values within each dataset and condition are shown in bold and underlined , respectively. AutoXRD uses a fixed candidate library (GanJiang dataset). PXRDGen-Flow’s reduced RRUFF and opXRD subsets satisfy its ≤100 -atom limit. PXRDGen and XtalNet abbreviate PXRDGen-Flow and XtalNet-HMOF100 (944 RRUFF and 654 opXRD).
XRD-only
XRD + composition
Model
Cov. ↑ (%)
Exact ↑ (%)
P ↑ (%)
R ↑ (%)
F1 ↑ (%)
Cov. ↑ (%)
Exact ↑ (%)
P ↑ (%)
R ↑ (%)
F1 ↑ (%)
MP500 ( N=1,000 )
AutoAnalyzer
100.00
1.70
46.49
58.08
51.12
95.63
23.37
85.08
58.08
69.25
S-CNN
100.00
1.27
40.11
18.99
25.06
43.17
1.60
42.46
18.99
25.72
AutoXRD
92.40
15.80
50.22
40.48
43.24
98.50
59.00
85.57
80.12
81.75
Gan Jiang
100.00
22.80
76.85
64.77
66.98
100.00
69.70
94.97
89.68
91.61
Table 2: Multiphase identification across simulated and experimental diffraction patterns. Each dataset contains 1,000 hidden mixtures. Results are reported under XRD-only and composition-assisted conditions. Higher values are better for all metrics; the best and second-best values within each dataset and condition are shown in bold and underlined , respectively. Cov. denotes the percentage of mixtures with at least one valid predicted phase after parsing, deduplication and any declared composition filtering. Exact denotes complete phase-set recovery. Exact, precision (P), recall (R), and F1 are evaluated over all 1,000 mixtures in each run. P, R, and F1 are sample-wise macro averages.
Track
Required prediction
Primary metrics
Single-phase identification
Ranked list of at most five CIFs
Top-1, Top-3, Top-5; MRR@5
Multi-phase identification
Unordered CIF set
coverage; exact set; macro precision, recall and F1
Refinement
Calculated intensity pattern
Rp , Rwp , correlation, score
Table S1: DeltaXRDbench task contract. A phase set is evaluated after one-to-one matching between submitted and reference structures.
Figure S1: Representative single- and multiphase samples in DeltaXRDbench. a,b , Simulated MP500 patterns; c,d , experimental RRUFF patterns; and e,f , experimental opXRD patterns. a,c,e show representative single-phase examples, pairing each normalized powder-XRD profile with its crystal structure. b,d,f show representative multiphase mixtures, with the combined profile at left and the constituent phase structures at right. The horizontal axis is 2θ (degrees) and the vertical axis is normalized intensity; labels identify the corresponding database records. Sphere colours distinguish atomic elements and dashed polyhedra indicate unit-cell boundaries.
Source
Single
Multi
Total
MP500 (simulated)
10,000
30,000
40,000
RRUFF (experimental)
1,164
10,000
11,164
opXRD (experimental)
880
10,000
10,880
All sources
12,044
50,000
62,044
Table S2: Dataset composition in the current DeltaXRDbench release. “Single” denotes single-phase patterns, whereas “Multi” denotes multiphase patterns generated by mixing single-phase patterns.
Quantity
Range / setting
Grain size
10–100
Preferred orientation (each component)
−0.3 – 0.3
Thermal vibration
0–2
Zero shift
0–2
Detector–sample distance
300–600
Detector half-height / sample half-height
3–8 / 1–4
Table S3: Default ranges sampled for dynamic Pysimxrd simulation.
Component
Setting
Input grid
3,500 samples; 10–90 ∘
FFT views
original + 3 low-pass views
CNN base channels
64
Attention dimension / heads
192 / 6
XRD tokens / element queries
96 / 4
Attention refinement blocks
2
Table S4: Default model configuration. The parameter count was computed from count_parameters(Xmodel()) .
Figure S2: XDecomposer learns to separate multiphase diffraction signals without a predefined phase list. a , Two-stage learning. Stage I pretrains a global encoder by reconstructing masked regions of single-phase patterns. Stage II combines a hierarchical encoder, the frozen global encoder (snowflake), a feature-wise linear modulation (FiLM) phase guide and a symmetric decoder to produce a set of candidate component patterns from a mixed XRD input. b , Inference framework. Hierarchical multi-scale features and global context are conditioned by learned phase queries through cross-attention and aggregation; the decoder outputs sigmoid masks mk , which are applied to the input profile x to yield physics-consistent component candidates y^k . Curves represent diffraction intensity as a function of 2θ ; colours distinguish predicted components.
Category
Parameter
Value or range
Random perturbations
Crystallite size (nm)
10–120
Thermal vibration coefficient
0.01–0.20
Zero shift
0–0.20
Detector–sample distance (mm)
300–600
Detector slit half-height (mm)
3–8
Sample half-height (mm)
1–4
Table S5: Parameters used to generate simulated pretraining data for XDecomposer.
Figure S3: XRDinspector provides transparent quality assessment for PXRD results. a , Candidate CIF(s) and an experimental pattern are aligned to quantify peak coverage, precision, phase support and intensity agreement. b , Observed and calculated profiles are evaluated using Rp , Rwp , correlation, overlap and local-residual diagnostics. c , The unified policy converts these diagnostic scores into four delivery levels, from high-confidence delivery (L1) to rejection (L4), and returns an auditable report with metrics, warnings and settings.
Component
Role
Experiment
Provides diffraction data, structural candidates and instrument metadata.
Gan Jiang
Audits inputs, defines hypotheses, evaluates outputs and reports conclusions.
Automation
Isolates each run, executes the backend and records logs.
FullProf
Calculates diffraction profiles and performs least-squares refinement.
Table S6: Roles in the Gan Jiang–FullProf refinement workflow.
Figure S4: Evidence-governed refinement workflow for the FullProf and GSAS-II branches. Experimental XRD data, structural CIFs or PCR files, and instrument records first undergo language-model preflight checks of links, units, phase model and radiation assumptions. Insufficient evidence triggers a request for calibration or phase-identification information. The FullProf branch (top) copies inputs into a new PCR file, performs refinement and either preserves the resulting profiles, logs and JSON summary or proceeds to the next minimal parameter-group test. The GSAS-II branch (bottom) validates its staged inputs, executes Rietveld refinement, and subjects the result to physical and statistical review before recording project files, metrics and tracebacks. Yellow chevrons denote the ordered refinement stages; diamonds denote decision gates. Accepted analyses retain diagnostics, limitations and a qualified conclusion.
Table S7: Programmatic access to selected capabilities of the Gan Jiang diffraction toolchain. Endpoint identifiers are supplied to the --endpoint argument of delta-cli science invoke --tool xrd .
Figure S5: Compositional and structural landscape of the MP500 dataset. a , Principal-component map of elemental composition for all 100,315 stored structures; colour denotes the number of structures per hexagonal bin. b , Co-occurrence network of the 16 most prevalent elements; links represent the 28 largest pairwise co-occurrence counts. c , Distribution of chemical complexity, expressed as the number of distinct elements per structure. d , Empirical cumulative distribution of atoms per stored unit cell; vertical lines mark the median and 90th percentile. e , Pairwise co-occurrence matrix of prevalent elements.
Figure S6: Overview of the experimental candidate library. a , Distribution of the number of atoms in the stored unit cells; solid outlines show binned fractions, dashed curves show empirical cumulative distributions, and dotted vertical lines mark the medians. b , Crystal-system distributions within RRUFF and the combined opXRD + private-data branch; symmetries were identified with spglib using symprec=10−2A˚ , and numbers above the bars give record counts. c , Occurrence of the 12 most prevalent elements, with each element counted at most once per structure record. d , Source composition of the collection, with distinct CIF counts determined from the stored SHA-256 checksums.
Figure S7: Five-dimensional index of the MP500 XRD archive. a , Archive-wide occupancy of 2,934,906 stored theoretical reflection positions, with a modal value at 30.9 ∘ 2 θ . b , Structure-resolved diffraction barcode for 50 fixed-seed records; colour denotes normalized reflection intensity. c , Rank-frequency distribution of leading space groups. d , Hexagonally binned relation between stored-peak count and Shannon entropy of normalized peak intensity. e , Elemental support of the inverted index; colour denotes the number of structures containing an element.
Phase
a (Å)
b (Å)
c (Å)
α ( ∘ )
β ( ∘ )
γ ( ∘ )
Gypsum
5.6800(9)
15.2144(0)
6.5303(7)
90
118.4841(5)
90
Phosgenite
8.1600(0)
8.1600(0)
8.8834(3)
90
90
90
Cerussite
5.1794(7)
8.4922(9)
6.1418(1)
90
90
90
Galena
5.9388(0)
5.9388(0)
5.9388(0)
90
90
90
Laurionite
9.7002(6)
4.0200(3)
7.1108(1)
90
90
90
Table S8: Refined lattice parameters of the five mineral phases identified in the ancient Egyptian cosmetic powder. Values in parentheses are the estimated uncertainties in the final reported digit.
Figure S8: Whole-pattern decomposition of binary and ternary powder mixtures. a , Fits to a nominal 2:8 SiOX2 ( α -quartz)– TiOX2 (rutile) mixture (top) and a nominal 1:1:1 AlX2OX3 (corundum)– ZnO (wurtzite)– NaCl (rock salt) mixture (bottom). Black curves show the measured scans, grey dashed curves show the fitted background and coloured fills show the phase-resolved contributions. b , Component diffraction profiles recovered for each phase in the two mixtures, using the same colour code.
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA) · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Guangzhou Municipal Key Laboratory of Materials Informatics, The Hong Kong University of Science and Technology (Guangzhou) +1