Data, Numbers, and Geometry: Three Tutorials on Numerical Methods, Machine Learning, and Evaluation
Authors: Jessica N. Howard, Yidi Qi, Tomás S. R. Silva
Organizations: Instituto de Matemática, Estatística e Computação Científica (IMECC), Universidade Estadual de Campinas (UNICAMP), Rua Sérgio Buarque de Holanda, 651, 13083-859 Campinas, São Paulo, Brazil
We present three practical tutorials on numerical computation and machine learning for mathematical research, developed for the DANGER: Data, Numbers, and Geometry workshop held at the Banff International Research Station in April 2026. The first develops a numerical approach to exterior calculus from pointwise evaluations of differential forms, using a flux formulation of the exterior derivative. Examples in Euclidean space and on the sphere illustrate geometric identities, topological features, and the effects of approximation and finite precision. The second examines how mathematical structure guides neural network design through examples involving elliptic curves, quivers, and a boundary value problem. It explores how architectural choices affect learning and uses interval arithmetic to bound the residual of a trained network over the full interval of the boundary value problem. The third addresses the evaluation and presentation of machine learning results, covering performance metrics, statistical uncertainty, classification thresholds, receiver operating characteristic curves, and accessible figure design. Throughout, the tutorials distinguish numerical agreement, predictive accuracy, structural guarantees, and rigorous bounds as different forms of evidence. Each contribution can be read independently, with accompanying notebooks and exercises that allow readers to reproduce the examples and adapt the methods to other problems.
Figures & tables
Figure 1 : Truncation vs. cancellation for the centered stencil of Eq. ( 3.1 ). The optimum ε∗≈10−5 balances the two regimes; under oracle noise of amplitude σ it shifts to ε∗∼σ1/3 .
0-form f
1-form ω
2-form η
output of d
1-form df
2-form dω
3-form dη
classical name
gradient
curl
divergence
probe region
segment [a,b]
parallelogram
parallelepiped
boundary ∂
2 oriented points
4 oriented edges
6 oriented faces
boundary integral
f(b)−f(a)
edges∑∫ω
faces∑∬η
d2=0 says…
\curl(\gradf)=0
\dive(\curlω)=0
(top degree)
Table 1 : The unified picture in R3 : one algorithm, three classical operators.
z0
surface integral
line integral
∣diff∣
0.0
6.2831853
6.2831853
∼10−8
0.5
4.7123890
4.7123890
∼10−8
0.9
1.1938052
1.1938052
∼10−8
Table 2 : Stokes’ theorem on geodesic caps Σ={z≥z0}⊂S2 , verified from black-box evaluations only. Exact value: 2π(1−z02) .
Data type
Key structure
Architecture(s)
Inductive bias
Tabular
unstructured features
MLP, boosted trees
smoothness; axis-aligned splits
Sequential
order, locality
1D CNN, transformer, SSM
shift equivariance; attention; recurrence
Grid / image
2D spatial structure
2D CNN, ViT, U-Net
2D translation equivariance, locality
Set / multiset
no ordering
DeepSets, set transformer
permutation-invariant aggregation
Graph
nodes, edges, topology
GNN (message passing)
permutation equivariance, locality
Geometric
symmetry group ( E(3) , …)
equivariant networks
group equivariance by construction
Table 1 : The landscape: from the structure of the data to the architecture that encodes it.
Figure 1 : (a) Class-averaged Frobenius traces ap over the training split. (b) The same after Hasse normalization, a~p=ap/(2p) . In both, the two class averages are near mirror images and their separation is largest at small p . (c) Two individual curves over the first 100 primes: the per-curve signal is indistinguishable from noise, so the pattern is a property of the sequence , recoverable only in aggregate.
Model
Input
Parameters
Train acc.
Test acc.
1D CNN
1,229 primes
39,714
99.91%
99.88%
MLP
1,229 primes
39,922
100.00%
99.13%
Transformer
100 primes
11,746
99.53%
99.51%
Table 2 : Rank classification on 3,468 held-out curves. The CNN and MLP have matched parameter budgets and see all 1,229 primes; the transformer sees only the first 100 ( p≤541 ), where attention’s O(T2) cost is affordable, so its accuracy is not directly comparable to the other two.
Figure 2 : Rescaled per-class saliency of the 1D CNN at three training checkpoints. At epoch 1 the map is unstructured noise. By epoch 4 it has become smooth and antisymmetric between the two classes, with oscillatory structure below p≈4000 . By epoch 7 the sensitivity is sharply concentrated in the small primes and decays towards zero for large p : the network has localized the discriminative signal in exactly the regime where the class averages of Fig. 1 separate.
Evaluation set
Vertices n
Accuracy
Training set
7,8 (seen)
100%
Test set
9 (unseen)
99%
Table 3 : Quiver mutation-type classification (type A versus type D ), after eleven epochs. The model is trained on 800 quivers with n=7,8 and evaluated on 100 quivers with n=9 , a size never seen in training. Independently of training, the node equivariance of Eq. ( 3.4 ) and the invariance of the graph-level logits both hold with maximum error exactly 0.0 .
Figure 3 : (a) The trained PINN against the exact solution sin(πx) . (b) Training loss. (c) The interval enclosure of the residual −uθ′′−f over 104 subintervals (band) against the pointwise residual at subinterval centres (line). The enclosure is tight; the residual oscillates within about 3×10−2 across the interior and rises steeply at both endpoints.
Figure 1 : Figure showing how a ROC curve can be constructed from a score distribution. Various threshold values are swept from right to left. For each threshold the TPR (fraction of class B scores falling above the threshold) and TNR (fraction of class A scores falling below the threshold) are measured. These (TPR, TNR) points are collected to form the ROC curve. The points on the ROC curve corresponding to the three thresholds are indicated.
Figure 2 : Example of how a plot can be used to efficiently tell a story. We have modified a figure from Ref. [ 12 ] to convey three key points about the result: (1) theoretical predictions agree closely with the empirical measurements, (2) the losses exhibit semi power-law behavior, (3) the two classification tasks follow quantitatively different trends.
Figure 3 : Figure showing a potential fix for a plot with too many lines . The primary goal of the plot is to show that “2D” methods perform worse than “2D ΔpT ” methods across all choices of N . However, the plot is trying to do too many pairwise comparisons at once. Since this performance gain is present for all methods (minOT, QminOT 1%, and QminOT 10%), we can simply choose a representative method (in this case QminOT 1%) and say that we see a similar trend in the other methods in the caption. If comparing the performance of the three methods is also important, one can show this in a separate plot. The data and methods for this example were taken from Ref. [ 13 ] .
Figure 4 : Figure showing two examples of residual plots . When you want to show agreement between two quantities (in this case binned data and the best fit function) but also want to show the overall shape, it can be helpful to use a residual plot. In the main plot, we show the fitted function and binned data to give an idea of the agreement and shape of the distribution. However, more information about the fit can be gained by looking at the difference or ratio of the two quantities. This can be shown in a residual plot below the main plot. In this case, it helps accentuate differences happening at the tails of the distribution which would be difficult to notice otherwise.
Department of Mathematics, Imperial College London, UK · London School of Geometry and Number Theory, UK · Department of Electrical and Computer Engineering, Rice University, USA +2