A Physics-Guided Transformer Framework for Electromigration Analysis in Multi-Segment Interconnects
Authors: Pavlos Stoikos, Anuj Pathania, George Floros
Organizations: Dept. of Electrical and Computer Engineering University of Thessaly, Greece · University of Amsterdam Amsterdam, The Netherlands · Dept. of Electronic and Electrical Engineering, Trinity College Dublin, Ireland
As technology scales to smaller nodes, increasing current densities make electromigration (EM) one of the dominant reliability challenges in on-chip interconnects. Accurate transient stress analysis is needed to identify wires susceptible to EM degradation, but applying physics-based solvers across many interconnects remains computationally expensive. This paper proposes a physics-guided transformer framework for fast EM stress prediction in multi-segment interconnect lines. The framework converts each line into geometry- and DC-aware segment tokens and uses transformer attention to capture line-level context. A lightweight query decoder then predicts stress at selected locations and time instants. The model is trained with an objective that combines normalized supervised regression, linewise relative-L2 loss, and physics-guided continuity and terminal-flux terms. Experiments on IBM power grid benchmarks show that the proposed model achieves relative-L2 error below 8% and reaches up to 2459.68× speedup compared with the matrix exponential~solver.
Figures & tables
Figure 1 : Architecture and training flow of the proposed physics-guided transformer framework.
Benchmark
Total lines
Train/all (%)
Max seg.
Mean seg.
Max current (Pa)
IBM1
4160
30
55
7
1.15×109
IBM2
1202
60
236
104
9.14×108
IBM3
17513
23
965
47
6.40×108
IBM4
26302
15
572
35
1.58×107
IBM5
6483
62
577
166
1.03×108
IBM6
23595
41
1010
69
4.06×108
Table I : IBM line-dataset statistics.
Benchmark
Train rel- L2 (%)
Test rel- L2 (%)
Test RMSE (Pa)
Test MAE (Pa)
IBM1
7.3
6.8
1.97×103
1.20×103
IBM2
6.1
7.0
7.15×102
2.38×102
IBM3
5.8
7.8
2.72×105
6.73×104
IBM4
5.2
5.7
6.29×102
2.00×102
IBM5
4.4
3.7
3.18×101
1.86×101
IBM6
3.8
4.2
1.90×104
7.31×103
Table II : Evaluation results for the IBM line datasets summarized in Table I .
Figure 2 : Predicted-versus-ground-truth stress comparison on representative IBM test cases.
Figure 3 : Stress evolution along a 110-segment interconnect for multiple time points.
Benchmark
Golden (s/line)
Proposed (s/line)
Speedup
IBM1
0.019
5.21×10−5
367.95x
IBM2
0.139
5.65×10−5
2459.68x
IBM3
0.350
8.61×10−4
406.24x
IBM4
0.222
3.11×10−4
712.86x
IBM5
0.224
1.33×10−4
1675.76x
IBM6
0.389
5.51×10−4
706.53x
Table III : Average runtime comparison over the 10 largest extracted lines of each IBM benchmark.
Predicting electromagnetic field (EMF) strength in urban environments is essential for cellular network planning but computationally expensive with physics-based simulators. We propose a multi-conditioned dense prediction framework that generates 500 500 EMF maps from building layout images and antenna configurations. Our architecture uses a High-Resolution Transformer (HRFormer) backbone with two complementary conditioning mechanisms: Feature-wise Linear Modulation (FiLM) injects scalar antenna parameters into all backbone stages, while cross-attention fuses 1-D radiation pattern tokens with spatial features at the deepest stage. We further introduce transmitter-relative spatial channels encoding distance, proximity, and bearing from the antenna, enabling coordinate-consistent test-time augmentation (TTA) that reduces test MAE by 6.3%. To address the prediction difficulty imbalance across EMF maps, we design a composite loss combining masked L1, multi-scale structural similarity (MS-SSIM), and a focal L1 term that upweights high-signal pixels, outperforming individual loss components in all metrics. Our best model achieves a test MAE of 0.0461, a 25.2% improvement over a plain UNet baseline and 31.8% over an HRFormer-only baseline.Do-
Do-Eon Kim, Dongryul Park, Seungyoung Ahn +3
Soongsil University, Seoul 06978, Republic of Korea · Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology, Daejeon 34051, Republic of Korea · Department of Intelligent Semiconductors, Soongsil University, Seoul 06978, Republic of Korea
Full-wave electromagnetic (EM) simulation enables accurate patch-antenna analysis but is computationally expensive for large-scale forward prediction and inverse design. We present a mesh-native, physics-augmented graph-learning framework that treats radiation-pattern prediction as signal reconstruction on an irregular surface mesh. For the forward problem, a GPS graph transformer is trained with Physics-Augmented Intermediate Supervision (PAIS), an auxiliary node-level objective that predicts complex surface currents, the physical intermediate linking geometry to radiation. PAIS improves multiple GNN backbones at no inference-time cost, while shuffled-current and non-physical controls show the gain comes from physical correspondence. Direction-conditioned decoding and a differentiable radiation-integral consistency loss further exploit this structure. On an 80,000-sample CST benchmark, GPS+PAIS reaches MSE 0.17 / PSNR 19.67, generalizes to a PCA split, and transfers zero-shot to canonical patches. For inverse design, surrogate-filtered diffusion beats nearest-neighbor retrieval by 32% relative MSE.
Transformer-based surrogate models are increasingly used to replace expensive first-principles simulation in engineering design. But conventional transformer architectures are often over parameterized for the small, low-dimensional datasets typical of engineering design spaces, where large simulation data is expensive to generate. Under these conditions, excess parameter capacity leads to overfitting rather than improved accuracy, while also incurring unnecessary memory and compute overhead. This motivates a shift towards architectures that focus on additional compute rather than additional learnable parameters. This paper presents a hardware-aware evaluation of three recursive transformer paradigms for surrogate thermo-mechanical analysis of advanced packages: a)Tiny Recursive Model, b) our proposed Depth Recursive transformer, c) and a simple recursive transformer. We systematically compare their predictive performance (Recall, Mean Reciprocal Rank), parameter count, computational complexity (FLOPs), providing practical design guidelines for selecting recursive transformer architectures under resource-constrained scenarios. We validate this principle on two low-dimensional engineering prediction tasks: 1) thermo-mechanical reliability analysis of advanced semiconductor packages, where stress and warpage from thermal cycling must be evaluated repeatedly across a design-of-experiments sweep under costly finite element analysis (FEA). 2) Laplace PDE iterative numerical solver for capacitance field. Overall, recursive weight-sharing transformers provide an effective and generalizable trade-off between prediction accuracy, parameter efficiency, and computational cost for small data engineering surrogate modeling.
Kart-leong Lim
Institute of Microelectronics (IME), Agency for Science, Technology and Research (A*STAR), Singapore