Tree-like branching structures are common in nature, from botanical trees to neurons, blood vessels and respiratory trees. Their branching shape often reflects function, making structural modelling central to understanding how these systems work. Because acquiring real-world 3D data is often expensive or infeasible, realistic generative models are valuable for simulation and data augmentation. Existing morphology-specific models either constrain how topology is generated or rely on hand-tuned, mechanistic procedures. Generic 3D graph generators, by contrast, do not exploit or enforce the structure of trees. We propose Autoregressive Frontier Expansion, a generative framework that constructs trees through an iterative expansion process, simulating the biological growth of real trees. At each step, a flow-matching model parameterised by an SO(2)-equivariant GNN expands the frontier by predicting whether each active branch bifurcates or terminates. We evaluate our method on cortical neurons and botanical trees in unconditional, class-conditioned, and morphology-guided generation. Across both domains, the generated morphologies agree closely with the reference distributions and, in conditional experiments, with the specified targets.
Figures & tables
Figure 1: Overview of a generation step. Starting from a partial tree Tℓ , we (i) expand the current frontier (a subset of leaves) by deciding for each frontier node whether it branches or terminates, yielding a tree Tℓ+1 ; then (ii) localise by predicting 3D parent-relative offsets for the previous frontier nodes, producing Tℓ+1 . The process repeats until all frontier nodes terminate.
Figure 2: Stepwise generation of a botanical tree
Figure 3: Unconditional neuronal morphology generation. Top: Five representative generations from our model. Bottom: Distributions of eight morphology statistics for the training reference, SemlaFlow, MorphoGen, and our model. Dashed lines mark training-reference medians; grey values give the normalised W1 to the training reference (lower is better). Within-tree distributions weight each morphology equally, and horizontal ranges are determined from the training reference.
Structure
Distributional agreement
Method
Params.
Valid tree (%) ↑
Morph. ΔMMD2↓
TMD ΔMMD2↓
Coverage ↑
Density → ref.
Mean norm. marginal W1↓
Real–real
—
100.0
0.0000
0.0000
0.971
0.976
0.039
SemlaFlow
22.3M
76.6
0.2521
0.1038
0.223
0.155
0.389
MorphoGen
32.6M
100.0
0.8436
0.2077
0.006
0.018
1.515
Ours ( F=64 )
1.57M
100.0
0.0557
0.0146
0.808
0.827
0.175
Ours ( F=128 )
5.69M
100.0
0.0471
0.0108
0.805
0.862
0.163
Table 1: Unconditional neuronal morphology generation. Distributional metrics on 1,167 samples. The real–real row compares two disjoint training samples of the same size. Mean marginal W1 averages the six rotation-invariant statistics defined in Appendix F . Default for our method is F=256 .
Figure 4: Class-conditioned neuronal morphology generation. (A) Class-wise morphology discrepancy for our model and SemlaFlow relative to the training trees. The values represent normalised 1-Wasserstein distance, averaged over five rotation-invariant statistics (lower is better). Whiskers show 95% bootstrap intervals, resampling the generated morphologies while holding the sampled training reference fixed. n is the number of generated morphologies. (B) Two selected generations from our model per class. (C) Two selected training-set references per class.
Figure 5: TMD-conditioned generation. Top: Each row shows a conditioning target, one generated sample, their path-based persistence diagrams, and within-tree distributions of branch length, bifurcation angle and branch order. Bottom: 5 independent samples conditioned on the same target.
Validity
Persistence diagrams
Within-tree distributions
Method
Valid tree (%) ↑
Path PD W1↓
Radial-root PD W1↓
Branch length W1 ( μ m) ↓
Sibling angle W1 ( ∘ ) ↓
Branch order W1↓
Reference
100.0
5.914
6.453
16.15
8.34
1.190
SemlaFlow
63.2
3.158
2.453
18.01
10.61
0.668
Ours
100.0
2.939
2.803
13.38
9.50
0.548
Table 2: TMD-conditioned generation. Each method generates one morphology for each of 1,167 held-out targets. Entries in the persistence diagrams and within-tree distributions are means over target–generation pairs. The former compares normalised persistence diagrams, and the latter compares distributions of within-tree attributes including branch length, sibling angle, and branch order. The reference row pairs each target with a different held-out reference using a fixed-seed random derangement, providing a no-conditioning baseline.
Figure 6: Generated botanical trees across depth caps. Three representative samples are shown for each maximum topological depth dmax∈{10,15,20} .
Regime
Method
Valid tree (%) ↑
Max. path length ↓
Branch length ↓
Bifurcation angle ↓
Partition asymmetry ↓
Mean ↓
Unconditional
SemlaFlow
3.9
1.280
0.245
0.297
0.822
0.661
Ours
100.0
0.443
0.167
0.177
0.157
0.236
TMD-conditioned
SemlaFlow
12.5
0.982
0.172
0.417
0.683
0.564
Ours
100.0
0.165
0.087
0.238
0.076
0.141
Table 3: Botanical-tree generation. Population-level discrepancies on botanical trees restricted to depth 10 (full dataset). Morphology entries are tree-balanced 1-Wasserstein distances normalised by the reference standard deviation W1/σref (lower is better). Mean averages the distances. Bold marks the better result per training regime.
Depth
Method
Max. path length ↓
Branch length ↓
Bifurcation angle ↓
Partition asymmetry ↓
Mean ↓
Time (s/tree)
D10
SemlaFlow
1.314
0.275
0.330
0.737
0.664
0.562
Ours
0.542
0.178
0.197
0.148
0.266
0.192
D15
SemlaFlow
2.347
0.441
0.272
1.871
1.233
3.714
Ours
0.639
0.227
0.200
0.139
0.301
0.390
D20
SemlaFlow
2.572
0.361
0.452
3.964
1.837
12.317
Ours
0.702
0.250
0.198
0.366
0.379
0.617
Table 4: Unconditional botanical-tree generation across depth caps. Columns are W1/σref discrepancies (lower is better), Mean averages the four discrepancies, and Time is seconds per tree. Bold marks the better result within each depth.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: (Left) A tree with root r . The nodes are embedded in R2 . Their coordinates induce Euclidean edge lengths, which define a distance-to-root filtration. (Middle) A barcode showing the resulting H0 persistence under this filtration, with the elder rule deciding which branch survives each merge. The bars are annotated by ( vd,vb ), the death and birth nodes respectively. Note that the barcode itself is unlabelled. (Right) The corresponding persistence diagram. Points close to the diagonal represent short-lived features, while points further away represent more persistent features.
TMD channel
Assigned target W1
Matched alternative W1
Assigned/matched ratio ↓
Assigned closer (%) ↑
Path PD
2.916
5.543
0.526 ± 0.015
91.1 ± 0.8%
Radial-root PD
2.775
5.666
0.490 ± 0.015
90.8 ± 0.8%
Appendix
Table 5: Repeated sampling under TMD-conditioning. Persistence-diagram distances to the assigned target and to root-degree-matched alternatives across five generations per test neuron. “assigned closer” is the percentage of matched comparisons in which the assigned target has the lower distance. Variability between samples is much lower than to arbitrary neurons showing the effect of conditioning.
Figure 8: Error by cell type. Class-conditioned generation errors for our model (top) and SemlaFlow (bottom). Entries are class-wise W1 normalised by the reference standard deviation (lower is better). Both models work well on the same cell types and ours clearly outperforms SemlaFlow, except for node count that is given to SemlaFlow as input and partition asymmetry where errors are generally low.
Figure 9: Unconditional generation for neurons. Reference morphologies, post-processed SemlaFlow generations, MorphoGen generations, and generations from our model are shown. Neurons are drawn root-centred and to scale within each method.
Figure 10: Class-conditioned neuron samples. Each row corresponds to one of the seven cortical pyramidal-cell classes; columns compare training examples, SemlaFlow generations, and generations from our model. The displayed samples are selected per source and class rather than treated as paired reconstructions.
Figure 11: Stochastic variation under TMD-conditioning. Each row shows one held-out conditioning target followed by five independent generations for that target. The generated samples preserve the target’s coarse dendritic organisation while varying in branch geometry and topology. Row C shows a more difficult condition for which the generations are visibly more compact than the target. B4 and E0 are unrealistic since they exhibit two apical dendrites instead of just one.
We introduce a novel mathematical framework for analyzing and generating complex tree-shaped 3D objects, such as botanical trees and plants, which deform both in their 3D geometry and branching structure. Unlike previous works, which either consider only the skeletal structure of tree-like objects or approximate their 3D geometry using branch thickness, the proposed framework accurately models both the 3D geometry of the tree branches and the way they are interconnected. In this paper, we first generalize the Square Root Normal Fields (SRNF) representation, originally proposed for the statistical analysis of genus-0 surfaces, to tree-shaped 3D objects. We then treat tree-shaped 3D objects as points on a novel Riemannian tree-shape space equipped with a novel Riemannian metric that measures the amount of surface bending and stretching, and structural changes one needs to apply to one 3D tree-shape to align it with another. This way, deformations become trajectories in this novel tree-shape space. We analyze the theoretical properties of this novel tree-shape space and the corresponding metric and develop algorithms for computing point-wise and branch-wise correspondences and geodesic paths between complex 3D trees. We finally show how to use these building blocks for (1) computing statistical summaries, \ie means and modes of variation, of collections of tree-shaped 3D objects, and (2) synthesizing novel tree-shaped 3D objects by sampling from probability distributions fitted to a population of tree-shaped 3D objects. We demonstrate the performance and utility of the proposed framework on real and synthetic plants and botanical trees and show that it significantly outperforms the state-of-the-art.
Tahmina Khanam, Hamid Laga, Mohammed Bennamoun +4
Murdoch University, WA, Australia. · University of Western Australia, WA, Australia. · Johns Hopkins University, USA.
Generating realistic and diverse graphs is a key problem in machine learning, with applications in molecular discovery, circuit design, cybersecurity, and beyond. However, current graph generative models remain limited by scalability and novelty. Diffusion-based methods often require costly full-adjacency operations and long denoising chains, while many autoregressive and hybrid models have at least quadratic complexity. In addition, these models often imitate training graphs rather than generalize beyond them. We propose a lightweight autoregressive framework to address these issues. It uses a structure-guided topological ordering to serialize graphs into regular edge sequences, enabling near log-linear generation, and a two-phase training strategy that combines exploration-oriented augmentation with iterative refinement to reduce overfitting and promote controlled novelty. Experiments on molecular and non-molecular benchmarks show that our approach improves novelty while preserving high validity and uniqueness. The framework also supports both LSTM and Mamba-style causal sequence backbones, with large-memory accelerators enabling longer graph-sequence experiments beyond typical GPU limits.
Three-dimensional (3D) molecule generation has been dominated by diffusion models, which achieve strong generation quality but typically require the molecular size to be specified a priori. Recent autoregressive approaches have substantially narrowed the performance gap while naturally supporting variable-length generation and conditioning on partial molecular context. However, balancing unconditional and context-conditioned generation remains challenging. We introduce KRONOS, a latent autoregressive diffusion framework that generates molecules in the latent space of a pre-trained autoencoder, jointly modeling molecular graph topology and geometry, while retaining the flexibility of autoregressive generation. We further introduce a mixed training strategy inspired by Fill-in-the Middle (FIM) paradigm, enabling both unconditional and fragment-conditioned molecular generation within a single left-to-right autoregressive model. Experiments on QM9 and GEOM-Drugs demonstrate that KRONOS achieves leading unconditional generation performance among autoregressive methods, while remaining competitive with diffusion models. Moreover, fragment-conditioned generation is achieved with negligible impact on unconditional generation performance, demonstrating that both generation paradigms can be supported within a single architecture.
Federico Ottomano, Gaopeng Ren, Yingzhen Li +2
Department of Chemistry Imperial College London · Department of Computing Imperial College London