Predicting metal-organic framework (MOF) structures from given building blocks requires recovering their positions and orientations in a periodic crystal. The spatial effects of rotation errors are geometry-dependent and anisotropic. The same angular error can produce different atomic displacements depending on block size, shape, and rotation axis. Angular error alone, without reference to the specific block geometry, therefore cannot fully describe the spatial consequences of a pose error. We introduce AnchorPose, a meso-grained pose generation framework that incorporates this geometric dependence into its generative representation. It represents each block through a small set of representative atoms, combines their local geometry with the current spatial state, and generates their coordinates with Bayesian Flow Networks. Known atom correspondences enable rigid alignment to recover complete building-block poses and return geometrically consistent points to the generation process. This design connects point-level spatial prediction with block-level structural constraints. Geometry participates in the pose state and its prediction, while rigid reconstruction preserves intra-block structure without treating all atomic coordinates as assembly variables. On the MOF benchmark, AnchorPose improves single-candidate match rates over the compared block-level and all-atom baselines.
Figures & tables
Figure 1: Overview of AnchorPose. (a) The same rotation angle can lead to very different spatial displacements for building blocks with different geometries. (b) AnchorPose represents each building block by a small set of fixed-identity anchor atoms, generates their coordinates with a Bayesian Flow Network conditioned on local geometry and crystal context, and recovers the full block pose via Kabsch alignment. The rigidly reconstructed anchors are fed back as the clean estimate for the next generation step, enabling geometry-aware and structurally consistent MOF assembly.
Figure 2: From geometry-dependent pose error to rigid recovery. Left. Reference and 20∘ -rotated poses show how atomic displacement depends on block shape, distance from the axis, and rotation axis. Blue markers show projected displacements. Right. Kabsch aligns local anchors with predicted points and rotates the complete block. Numbers identify corresponding atoms. Translation is omitted.
Figure 3: Pose error and conformation residual. Rigid alignment removes pose error in (a) , but leaves the internal twist in (b) . Dark atoms show the rigid fit. Blue outlines show its perturbed or crystal counterpart.
Figure 4: Anchor prediction and rigid recovery. Local geometry, noisy anchors, and block context condition prediction. Alignment supplies Rt and recovers rigid blocks. Teal marks one block, black the others, and orange the anchors. Periodic images complete boundary-crossing blocks.
Loose ( stol=0.5 )
Strict ( stol=0.3 )
Model
Block geom.
Granularity
Rigid
Samples
MR ↑
RMS ↓
MR ↑
RMS ↓
1
0.23
0.3896
0.01
0.1554
DiffCSP
No
Fine
No
5
0.87
0.3982
0.08
0.1299
1
21.93
0.3329
5.28
0.2036
MOFFlow
Yes
Coarse
Yes
5
32.71
0.3290
8.68
0.2039
1
27.66
0.3234
6.96
0.1979
Table 1: Generation granularity and structure recovery. Block geom. denotes intra-block geometry. MR is in %. Baseline results follow AtomMOF ( Kim et al., 2026 ) .
Variant
Anchor
Anchor
Geom. rot.
Block
Pose
Loose
Strict
state
pred.
noise
fitting
updates
MR
MR
Anchor-state variants
AnchorPose
✓
✓
–
✓
✓
39.80
12.36
Direct rotation prediction
✓
–
✓
✓
38.89
11.89
Rotation-state alternatives
Direct rotation prediction
✓
✓
36.41
9.86
Table 2: Pose representation and component ablations. All variants use actual block geometry. Without anchor states or predictions, models use Bingham states or rotation outputs, respectively. Checkmarks indicate enabled components. MR (%) covers 19,792 structures with one candidate.
Anchors
4
6
8
Loose MR
39.80
39.39
38.82
Strict MR
12.36
12.31
12.00
Table 3: Single-candidate MR (%).
Figure 5: Pose errors under the native Gaussian-anchor and Bingham schedules. Medians over the paired validation probe. Panels (a,b) show input and recovered-pose errors over time. Panels (c,d) show step-50 input errors by block size. Spatial errors measure full-template RMS displacement from the target rigid pose. Size groups are tertiles of the full-atom RMS radius. Recovered poses include Kabsch alignment and inter-block updates. Appendix C details the probe protocol.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Noise acts on different pose variables. From the same clean block, (a) a noisy Bingham orientation moves the template rigidly, whereas (b) Gaussian anchors move independently. Gray denotes the clean reference. Panel (b) undoes scale and time-dependent shrinkage without rigid projection. Arrows show one native corruption draw.
Selection
Median
90th percentile
95th percentile
FPS
1.66
3.06
10.34
Random four atoms
3.80
16.59
29.01
Appendix
Table 4: Directional geometric coverage of four anchors. Distortion D is scale-normalized. Lower is better. Statistics count block occurrences across the full test set.
Setting
Value
Local-fitting network
Width 384, four atom-message layers, simultaneous torsion head, four coordinate-residual layers
Local-fitting loss
Circular torsion loss and aligned linker squared error, each with weight 1
200 epochs, 260 updates/epoch, batch 768, seed 348, validation-selected checkpoint at epoch index 199
Assembly atom / block GNN
Four atom layers, width 64, six block layers, width 512, time embedding width 128
Anchor decoder
Up to four FPS atoms, four Transformer layers, width 128, coordinate scale s=10 Å
Appendix
Table 5: Training and architecture settings. Batches are effective global batches. Both stages use BF16 mixed precision and AdamW.
Stage
Parameters (M)
Time / batch (s)
Peak memory (GiB)
Local fitting
8.34
0.059
0.92
Assembly (50 steps)
11.99
1.181
0.28
Appendix
Table 6: Model size and inference cost. Mean time per batch of 32 structures. Peak allocated memory over the profiled batches.
Figure 7: Local conformation preparation. Shared atom features predict torsions, followed by equivariant coordinate refinement. Assembly uses the resulting fixed block.
Training with one sampled time per example
1. Sample k and draw lattice, center, and anchor states from their clean targets at t=(k−1)/50 .
2. Encode atoms and blocks. Predict the lattice, centers, and raw clean anchors.
3. Recover rigid poses using Kabsch alignment and two inter-block point-based updates.
4. Reconstruct Pproj=RA and optimize Equation 29 .
Generation with fifty decoder calls
1. Initialize lattice, periodic-center, and anchor states at their priors.
Appendix
Table 7: Training and generation with fixed fitted blocks.
Population
Count
RDKit RMSD
Fitted RMSD
Reduction
All blocks
123,861
0.842
0.608
27.86%
All linkers
87,150
1.150
0.817
29.00%
Flexible linkers
76,676
1.292
0.914
29.21%
Appendix
Table 8: Local fitting reduces rigid-aligned error on held-out blocks. RMSD is in Å. Each row averages over the listed population.
Supervision before pose updates
Loose MR
Strict MR
Rotation loss
35.96
9.64
Anchor-coordinate loss
36.41
9.86
Appendix
Table 9: Rotation-output models with explicit geometry. Both use Bingham states, fitted blocks, actual point inputs, and final coordinate supervision. Match rates (%) cover all 19,792 test structures.
Step
State / point decoder
Input
Kabsch estimate
Recovered pose
20
Gaussian anchors
2.047
0.935
0.881
20
Bingham rendered points
0.957
0.515
0.438
30
Gaussian anchors
1.182
0.555
0.517
30
Bingham rendered points
0.786
0.429
0.361
40
Gaussian anchors
0.635
0.335
0.302
40
Bingham rendered points
0.674
0.399
0.331
Appendix
Table 10: Spatial pose error from input to recovered pose. Median full-template RMS displacement in Å, relative to the same target rigid pose.
Gaussian anchors
Bingham points
Anchor balance
Input
Final
Input
Final
Final difference †
Low
17.80
10.80
13.76
9.87
−0.93
Medium
11.29
6.54
13.48
9.46
+2.92
High
14.90
6.38
13.49
12.28
+5.89
† Bingham minus Gaussian.
Appendix
Table 11: Step-50 angular error by anchor balance. Means in degrees. Bingham has lower mean final error in the low-balance group. Gaussian anchors have lower means in the other two groups.
Variant
Loose MR
Strict MR
State process with actual points and a relative rotation head
Bingham state, rigidly rendered points
32.185
7.801
Gaussian state, rigidly projected points
29.815
6.796
Anchor attention with a Gaussian state and point head
Within-block anchor attention
36.116
10.459
MOF-wide anchor attention
36.045
10.272
Appendix
Table 12: Additional architecture and supervision controls. Each group is a separate paired comparison, using final 500-epoch EMA weights, 50 steps, one candidate, and all 19,792 test structures.
Figure 8: Three modeling resolutions for the same block. Block pose, selected corresponding atoms, and all atomic coordinates define different generation variables. AnchorPose predicts the orange-ringed atoms and reconstructs the complete block by rigid alignment. Panels share camera and scale.
Figure 9: Accurate linker orientation in a generated structure. Predicted and target rotations place the same 58-atom fitted template. Blue shows the prediction and dashed gray shows the target.
Figure 10: A pose error in the same generated structure. The accompanying 86-atom linker illustrates a failure under the same fixed-correspondence, pose-only evaluation.
Figure 11: Local fitting reduces a large conformational discrepancy. We rigidly align initial, fitted, and crystal coordinates to a common reference. Angular diagrams show measured dihedrals.
Figure 12: A failure of local fitting. Rigid assembly alone cannot correct an internal discrepancy remaining in the fitted template.