Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.
Figures & tables
Figure 2: Reach-budget plots for MNIST INRs. Fraction of 1,000 held-out teachers whose students reach a prescribed performance threshold defined relative to the teacher as a function of distillation budget (log scale), comparing CrossGMN with its anchor and other baselines (Appendix D.3 ). Each compression setting includes two panels: (a) Fidelity to the teacher , measured by PSNR(student,teacher)≥T dB, and (b) Student performance retention , measured by PSNR(student,image)≥PSNR(teacher,image)−ε . The thresholds T and ε are chosen separately for each compression setting to account for differences in student architecture capacity. See Section 5.1 for the full description of the reach-budget plots.
Table 2
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Mapping a teacher to an optimal student can be discontinuous. The width- 2 teacher parameters vT(t) in Equation 14 vary continuously and realize ft=(1−t)u+(1+t)v . By Lemma 3 , the optimal width- 1 student realizes (1−t)u for t<0 and (1+t)v for t>0 ; at t=0 , both u and v are optimal. These one-sided limits differ, so no continuous parameter map can select an optimal student for every teacher along the path (Proposition 4 ).
Family / operator group
Symmetry
dhid
Layers
Agg.
dv/de
∣θ∣
s
MNIST SIREN (all; depth and width+depth use s=.01 )
sign
128
4
mean
9/9
2,647,895
.005
ModelNet40 depth 4→2
sign
192
4
add
9/9
5,937,947
.02
ModelNet40 depth 4→1
sign
192
4
add
9/9
5,937,907
.02
ModelNet40 width
sign
192
4
add
9/9
5,937,907
.005
ModelNet40 width+depth
permutation
324
4
add
9/9
6,010,354
.01
ViT depth
scale
64
4
mean
25/17
405,587
.1
Appendix
Table 8: CrossGMN configurations. dv/de are node/edge input widths and ∣θ∣ is the metanetwork parameter count; s is the residual scale in Equation 5 .
ViT
CNN
MLP
MNIST INR
ModelNet40 INR
Objective
Eq. 27
Eq. 27
Eq. 27
Eq. 28
Eq. 29
Optimizer / decay
AdamW / 10−3
AdamW / 10−3
AdamW / 10−3
AdamW / 10−3
AdamW / 10−3
Learning rate
3⋅10−4 or 10−3
10−4
10−4 / 10−3
5⋅10−4 – 2⋅10−3
3 – 5⋅10−4
Warmup (steps)
none
200
200
none
none
Functional inputs
128 images
128 images
128 images
784 coordinates
4096 or 2048 points
Appendix
Table 9: Operator training.
Family
Reduction axis
Anchor
MNIST SIREN
width / depth / width+depth
magnitude pruning / layer fold / fold, then prune
ModelNet40 SIREN
width / depth / width+depth
magnitude pruning / strided layer transfer / strided prune
CIFAR-10 MLP
width / depth or width+depth
magnitude pruning / layer fold (scale 0.1 )
Heterogeneous MLPs
all
layer fold (scale 0.1 )
CIFAR-10 ViT
attention dimension / head count / depth / head count+depth
top- k attention dimensions / head drop / strided block transfer / strided blocks + head drop
CNN
channel width
magnitude pruning
Appendix
Table 10: Anchor initializations used for the distillation comparisons.
ViT
CNN
MLP
MNIST INR
ModelNet40 INR
Optimizer
AdamW
AdamW
AdamW
Adam
Adam
Learning rate
10−3
10−3
10−3
10−2
10−2
Batch
256 images
256 images
256 images
full 28×28 grid
4096 points
Augmentation
none
none
crop + flip
–
–
Cohort
1,000
400
1,000 (500 OOD)
1,000
1,000
Appendix
Table 11: Evaluation-time distillation. Classifiers use pure KL at temperature 1 and no learning-rate schedule; all evaluation loops record every 10 steps.
Reduction axis
Student
Params
Compression
Anchor
Operator lr
s
Width
[2,16,16,1]
337
3.52×
magnitude pruning
10−3
.005
Width
[2,8,8,1]
105
11.29×
magnitude pruning
2⋅10−3
.005
Width
[2,4,4,1]
37
32.03×
magnitude pruning
10−3
.005
Depth
[2,32,1]
129
9.19×
layer fold
5⋅10−4
.01
Width+depth
[2,16,1]
65
18.23×
fold, then prune
5⋅10−4
.01
Appendix
Table 12: MNIST SIREN configurations. Compression is the teacher-to-student parameter-count ratio; s is the CrossGMN residual scale.
Table 15: ViT compression configurations. All four settings use the strided cross-edge policy of Definition 5 ; when the teacher and student block counts agree it reduces to shared-index matching. Here d is the residual width, L the number of blocks, H the number of heads, and da the total attention dimension across heads.
Figure 4: MNIST INR reach–budget plots ( N=1,000 per setting). All five reductions in Table 1 ; fidelity reach above and performance-retention reach below.
Figure 5: ModelNet40 INR reach–budget plots ( N=1,000 per setting). All four reductions in Table 1 , showing fidelity to the teacher field. The asterisk marks a fallback threshold.
Figure 6: CIFAR-10 MLP reach–budget plots ( N=1,000 per setting). The three reductions in Table 2 ; fidelity reach above and performance-retention reach below.
Figure 7: CIFAR-10 ViT reach–budget plots ( N=1,000 per setting). The four reductions in Table 2 ; fidelity reach above and performance-retention reach below.
Figure 8: Transfer across datasets ( N=1,000 per setting). MNIST-trained CrossGMN operators applied to FashionMNIST INRs without retraining (Table 5 ). Asterisks mark fallback thresholds.
Figure 9: Reach–budget plots for CNNs and heterogeneous MLPs. (a) CNN width reduction ( N=400 , Table 2 ). (b) Four teacher architectures mapped to [3072,8,10] ( N=1,000 , Table 3 ). (c) The same operator on unseen source architectures ( N=500 , Table 5 ).
Setting
Anchor
Budget
Median speedup
IQR
>1 (%)
>2 (%)
Failed (%)
Wall-clock speedup
\Delta@ Budget
MNIST SIREN INRs
width 3.5 ×
magnitude pruning
2,710
2.20×
[1.24,3.86]
83
54
2
2.17×
+1.78
width 11.3 ×
magnitude pruning
2,480
5.20×
[1.14,12.48]
76
68
18
4.99×
+1.65
width 32 ×
magnitude pruning
3,520
4.10×
[0.00,14.10]
62
58
35
3.98×
+0.47
depth 2 → 1
layer fold
4,560
4.07×
[2.21,6.57]
90
78
7
4.00×
+0.65
width+depth 18.2 ×
fold, then prune
2,750
8.89×
[3.56,16.83]
86
83
13
8.36×
+0.44
Appendix
Table 17: Complete compression results. Budget is B from Equation 37 . Speedup statistics include failures ( SvT=0 ). Wall-clock speedup is the median per-teacher ratio from Equation 38 . \Delta@Budget is the median paired CrossGMN-minus-anchor quality difference at the last evaluation step not exceeding B : PSNR (dB) for INRs and test agreement for classifiers.
Teacher population
Teachers
Agreement ↓
Error IoU ↓
κ↓
CIFAR-10 MLP
1,000
0.715
0.694
0.682
CIFAR-10 ViT
1,000
0.792
0.556
0.769
MNIST CNN
400
0.891
0.210
0.879
CIFAR-10 MLP, four source architectures
1,000
0.643
0.683
0.602
unseen source architectures
500
0.618
0.653
0.574
Appendix
Table 18: Classifier teacher-population diversity. Mean statistics over 10,000 pairs of distinct teachers. Lower agreement, error IoU, and κ indicate greater diversity in their respective aspects.
Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.
Metanetworks are neural architectures designed to operate directly on pretrained weights to perform downstream tasks. However, the parameter space serves only as a proxy for the underlying function class, and the parameter-function mapping is inherently non-injective: distinct parameter configurations may yield identical input-output behaviors. As a result, metanetworks that rely solely on raw parameters risk overlooking the intrinsic symmetries of the architecture. Reasoning about functional identity is therefore essential for effective metanetwork design, motivating the development of equivariant metanetworks, which incorporate equivariance principles to respect architectural symmetries. Existing approaches, however, typically enforce strict equivariance, which imposes rigid constraints and often leads to sparse and less expressive models. To address this limitation, we introduce the novel concept of quasi-equivariance, which allows metanetworks to move beyond the rigidity of strict equivariance while still preserving functional identity. We lay down a principled basis for this framework and demonstrate its broad applicability across diverse neural architectures, including feedforward, convolutional, and transformer networks. Through empirical evaluation, we show that quasi-equivariant metanetworks achieve good trade-offs between symmetry preservation and representational expressivity. These findings advance the theoretical understanding of weight-space learning and provide a principled foundation for the design of more expressive and functionally robust metanetworks.
Viet-Hoang Tran, An Nguyen, Benoît Guérand +2
Department of Mathematics National University of Singapore
Neural network weights are increasingly a bottleneck for deployment, yet most compression pipelines treat layers independently and overlook cross-layer redundancy induced by function-preserving symmetries. We propose Motion-Compensated Weight Compression (MCWC), a weight-only codec that aligns permutation-symmetric blocks (e.g., hidden units and attention heads) to maximize cross-layer correspondence, turning depth into a predictable sequence. In the aligned coordinate system, MCWC uses a lightweight layer-sequential predictor with periodic keyframes and encodes only quantized prediction residuals using a learned entropy model trained under a rate distortion objective. A simple decoder reconstructs deployable weights by entropy decoding, dequantization, predictor-driven reconstruction, and inverse alignment, enabling fast weight materialization for inference. Across Transformer language modeling and vision classification, MCWC improves the rate accuracy Pareto frontier over strong quantization and learned weight-codec baselines, while maintaining competitive decode time. Ablations confirm that alignment, prediction, entropy modeling, and keyframe scheduling are each necessary for the full gains. Our code is available via https://github.com/Ism-ail11/MCWC.
Ismail Lamaakal
Multidisciplinary Faculty of Nador · Mohammed Premier University · Oujda 60000, Morocco