Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the application of tokens to physical simulation remains nascent. Because existing tokenizers are designed for the perceptual requirements of natural images, they struggle with scientific data, which exhibits large dynamic ranges and requires exact preservation of physical and spectral properties. In this work, we investigate the performance of a suite of image tokenizers across metrics designed to measure PDE fidelity. Observing that these baselines struggle to simultaneously capture fine geometric details and precise physical magnitudes, we propose Phaedra, a novel tokenizer inspired by classical shape-gain quantization and the paradigm of basis functions coupled with continuous coefficients. Phaedra acts as a highly effective nonlinear compression algorithm, massively reducing dataset footprints while maintaining physical fidelity. We demonstrate that Phaedra consistently improves reconstruction across diverse 2D gridded PDE solutions, generalizes robustly to unseen PDE types and real-world Earth observation data, and is competitive with continuous models in downstream proof-of-concept operator learning and masked autoencoding tasks.
Figures & tables
Figure 1 : Top Left: Density field of compressible Euler equations zoomed in to show fine details, ground truth and reconstructions with relative L1 errors. FSQ fails to capture high frequency information, resulting in smoothing of structures. Cosmos is designed to focus on fine details, but fails to capture precise amplitudes. Phaedra is able to model both phenomena accurately, minimizing reconstruction errors. Bottom Left: Phaedra tokenization pipeline. Embeddings are split and encoded in two streams: 1-dimensional amplitude tokens (finely quantized with 1024 levels) and Cμ -dimensional morphological tokens (quantized via multi-dimensional FSQ). Right: Visualization of the disentangled representation. The morphological component, generated here as a reconstruction with amplitude tokens set to zero, encodes local structure and high frequency features. The amplitude tokens encode a globally coherent, smoothed representation of the signal. Learning global and local features separately improves the final quality of the reconstruction.
Figure 2 : Even after normalization, natural images have a much narrower range than scientific data.
Method
K
Φ
Q
Ψ
# (Tokens)
VQ-VAE
1
Id.
VQ
Id.
hw
FSQ
1
Id.
FSQ
Id.
hw
VQ-VAE-2
2
Scale Split
VQ × 2
Con.+Ups.
h1w1+h2w2
VAR
S
Residual
FSQ
Add
∑shsws
Phaedra
2
Channel Split
FSQ × 2
Learn
2hw
Table 1 : Structural comparison of tokenization schemes.
Table 4Figure 5Table 6
Figure 5 : Downstream Operator Learning. A downstream transformer is trained using cross-entropy loss to predict the dynamics of the compressible Euler equations with the Riemann-curved (RC) initial conditions. The results show the prediction at the final timestep, using only the initial conditions as inputs. The amplitude tokens clearly capture much of the overall dynamics in density (top) and u− velocity (bottom), while fine structure is added via morphology token prediction. Additional figures for all downstream tasks are available in SM.6 .
Figure 6 : An example of an input, its amplitude tokens, and reconstructions from the (i) morphological tokens with amplitude tokens set such that their embeddings are equal to zero and (ii) full set of tokens.
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Dataset
nMAE ↓
nRMSE ↓
Δσloc2↓
γmin↑
Continuous
ID
0.672
1.122
2.98
98.4%
OD 1
1.154
2.079
3.40
99.2%
OD 2
1.967
2.821
3.02
97.0%
Phaedra
ID
1.522
2.489
5.96
93.6%
OD 1
1.224
2.435
6.47
98.1%
OD 2
3.147
4.237
5.83
77.6%
Appendix
Table 7 : Comparison of results for the continuous autoencoder compared to the discrete tokenizer Phaedra . ID denotes the test split of the training dataset. OD 1 denotes out-of-distribution datasets which are still defined by the Euler equations. OD 2 denotes datasets which are governed by PDEs not present in the training data. For the same encoder-decoder family and comparable training, the continuous autoencoder is expected to be at least as expressive because it does not have the quantization constraint; in our experiments, it gives the best reconstruction accuracy.
Figure 7 : Plots of the encoder outputs, following along from the embeddings, through the segregation of morphological and amplitude components, the quantization step, and their eventual recombination before decoding.
Figure 8 : Plots of the encoder outputs, following along from the embeddings, through the segregation of morphological and amplitude components, the quantization step, and their eventual recombination before decoding.
Figure 9 : Token swapping of the amplitude and morphology for the Phaedra image (A) and a density field from the CEU RC dataset (B). The reconstruction using amplitude tokens from A with morphology tokens from B maintains overall pixel values consistent with image A, but introduces small-scale vortices and flow patterns from B.
Ablation
Low-freq retention
High-freq retention
Zero Amplitude
3 – 5%
∼300% (artifacts)
Zero Morphological
103 – 106%
47 – 49%
Appendix
Table 8 : Frequency retention after targeted latent space ablation.
Figure 10 : A unified view of discrete tokenization. All methods share the encoder E and decoder D , and differ only in the factorization Φ , the quantizers {Qk} , and the recombination Ψ of Eq. 1 . Prior tokenizers either do not factorize at all (VQ-VAE, FSQ) or factorize across scale (VQ-VAE-2, VAR); Phaedra factorizes across channel , into a vector-quantized morphology stream and a densely quantized scalar amplitude stream, recombined by a learned channel mixing. Entries correspond to Table 1 .
Model
Parameters
Channel Mult.
Codebook Size
Token Count 128
Continuous AE
97M
[2,2,4]
-
-
FSQ
97M
[2,2,4]
8640
1024
IBQ
97M
[2,2,4]
16384
1024
VQ-VAE2
97M
[2,2,4]
20480
1280
VAR small
97M
[2,2,4]
8640
1704
VAR large
112M
[2,2,4]
8640
2728
Appendix
Table 9 : Model & quantizer configurations for the different models. The Codebook Size is the number of learnable (for IBQ, VQVAE2) or product of the levels (for FSQ) codes in the codebook. All models accept 128×128 inputs. The Token Count 128 is the number of tokens, assuming a 128×128 input.
Parameter
Value
Optimizer
AdEMAMix
Base Learning Rate
10−4
Optimizer Betas
( β1 =0.5, β2 =0.9, β3 =0.99)
AdEMAMix α
2.0
Weight Decay
0.01
EMA Decay
0.999
Appendix
Table 10 : Training hyperparameters across all models. We fix the autoencoder structure across all models, unless otherwise specified. All models are trained with the following hyperparameters.
Dataset Name
Initial Conditions
Vars
Train/Val/Test
Steps/Traj
CEU Gauss
Gaussian density perturbations
4
9,640/120/240
21
CEU KH
Kelvin–Helmholtz instability
4
9,640/120/240
21
CEU RC
Curved interface Riemann
4
9,640/120/240
21
CEU Riemann
4-Quadrant Riemann interaction
4
9,640/120/240
21
INS Gauss
Gaussian vortex field
2
19,640/120/240
21
INS Sine
Sinusoidal perturbations
2
19,640/120/240
21
Appendix
Table 11 : Summary of the scientific datasets used for training. “Vars" indicates the number of physical fields present in the simulation. Note that the standard ( 1282 ) and high-res ( 5122 ) runs use different variable bases for the Compressible Euler datasets (Primitive vs. Conservative).
Model
Dataset
ρ
u
v
p
Continuous AE
CEU Gauss
0.4413
0.3192
0.2958
0.3928
CEU KH
0.6495
0.3481
0.4031
0.5411
CEU RC
2.3518
1.5309
1.5208
1.4503
CEU Riemann
0.4697
0.6683
0.5064
0.3952
INS Gauss
—
0.2655
0.2679
—
INS Sines
—
0.2888
0.3274
—
Appendix
Table 12: Comparison of nMAE ↓ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
Continuous AE
CEU Gauss
0.6654
0.4270
0.4450
0.6197
CEU KH
1.3745
0.5130
0.7470
0.8017
CEU RC
4.8338
2.5375
2.5300
2.6390
CEU Riemann
0.9652
0.9132
0.8846
0.7325
INS Gauss
—
0.3753
0.3591
—
INS Sines
—
0.4071
0.4645
—
Appendix
Table 13: Comparison of nRMSE ↓ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
Continuous AE
CEU Gauss
10.81
4.94
5.09
11.11
CEU KH
29.07
6.56
10.57
7.40
CEU RC
80.56
34.85
34.90
44.92
CEU Riemann
22.18
16.26
16.02
16.22
INS Gauss
—
2.00
2.00
—
INS Sines
—
4.22
4.52
—
Appendix
Table 14: Comparison of normalized maximum error (n L∞↓ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
AE Continuous
CEU Gauss
7.1022
1.4656
1.4148
6.4110
CEU KH
2.5600
1.5330
6.2576
6.4077
CEU RC
2.4214
2.2839
2.1390
2.2797
CEU Riemann
2.0933
5.9525
3.1555
1.7560
INS Gauss
—
1.2037
1.2210
—
INS Sines
—
0.9001
1.0462
—
Appendix
Table 15: Comparison of local variance error Δσloc2↓ across all 4×4 models, datasets, and variables.
Model
Dataset
ρ
u
v
p
Continuous AE
CEU Gauss
99.99
99.83
99.82
99.99
CEU KH
99.98
99.86
93.92
100.00
CEU RC
96.65
89.54
90.00
99.91
CEU Riemann
99.88
99.00
99.04
99.89
INS Gauss
—
99.83
99.85
—
INS Sines
—
99.70
99.67
—
Appendix
Table 16: Comparison of γmin↑ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
Continuous AE
CEU Gauss
99.5375
98.6984
98.6770
99.5449
CEU KH
99.1300
97.3549
89.8986
99.7914
CEU RC
97.7406
96.5632
96.6366
98.1140
CEU Riemann
99.1136
98.7110
98.7952
99.3907
INS Gauss
—
98.8140
98.7005
—
INS Sines
—
98.1300
98.0971
—
Appendix
Table 17: Comparison of Log spectral energy fidelity Flog↑ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
Continuous AE
CEU Gauss
1.7245
3.1961
3.1904
1.8379
CEU KH
2.7869
3.3842
4.3291
0.5914
CEU RC
3.4748
3.7626
3.7706
3.1601
CEU Riemann
2.8192
3.3568
3.3694
2.5334
INS Gauss
—
6.1458
6.2652
—
INS Sines
—
6.3145
6.3756
—
Appendix
Table 18: Comparison of maximum spectral difference ΔPmax across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
FSQ
CEU Gauss
94.3750
89.1551
89.3403
94.2361
CEU KH
93.4259
87.8704
92.8356
94.7222
CEU RC
98.4838
99.0162
98.9815
98.3681
CEU Riemann
96.0301
94.2477
94.2245
92.9398
INS Gauss
—
53.7500
54.4792
—
INS Sines
—
86.7014
88.2523
—
Appendix
Table 19: Comparison of codebook utilization U↑ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
FSQ
CEU Gauss
10.4052
9.9868
9.9988
10.3707
CEU KH
9.7379
9.8794
10.5237
11.5640
CEU RC
12.0170
12.0983
12.1081
11.9914
CEU Riemann
10.3948
10.1950
10.2560
10.0298
INS Gauss
—
8.7338
8.7793
—
INS Sines
—
10.4527
10.5076
—
Appendix
Table 20: Comparison of token entropy H↑ across all 4×4 models, training datasets, and variables.
Model
Dataset
ρ
u
v
p
FSQ
CEU Gauss
20.4305
23.6296
23.5377
20.6941
CEU KH
25.5333
24.4507
19.5239
11.5686
CEU RC
8.1044
7.4829
7.4079
8.3005
CEU Riemann
20.5100
22.0374
21.5712
23.3010
INS Gauss
—
33.2113
32.8633
—
INS Sines
—
20.0666
19.6474
—
Appendix
Table 21: Comparison of token redundancy R↓ across all 4×4 models, training datasets, and variables.
Model
Dataset
nMAE ↓
nRMSE ↓
Δσloc2↓
γmin↑
Utilization ↑
Continuous AE
CEU RKH, ρ
2.6836
4.5891
5.1315
0.9788
—
CEU AIR, ρ
0.4410
1.1844
4.0184
0.9998
—
INS SVS, u
0.3382
0.4635
1.0631
0.9985
—
POI, u
0.7140
1.0170
1.27e-08
0.9671
—
DAR, u
0.4245
0.6117
1.6570
0.9981
—
ALC, u
1.4171
1.9240
3.6923
0.9894
—
Appendix
Table 22 : Results for out of distribution datasets. CEU RKH, AIR, and INS SVS comprise OD 1 , while all other datasets comprise OD 2 .
Figure 11 : Scaling with respect to bottleneck resolution. We observe that the model scales with respect to the absolute number of tokens as opposed to the ratio of downsampling. That is, Phaedra performs equally well using 162 downsampling on high resolution ( 5122 ) data as when using 42 downsampling on low resolution ( 1282 ) data, as each compresses the input to 32×32 tokens. We also observe a major drop-off in performance when using fewer than 32×32 tokens, with diminishing returns as the number of tokens increases. This is illustrated above for the normalized MAE, MSE, and minimal spectral coherence.
Figure 12 : Scaling with respect to amplitude codebook size. We trained Phaedra and FSQ on the CEU RC ρ dataset to observe scaling performance as the number of available amplitude tokens increases. While increasing the size of the amplitude codebook consistently yields better results, these gains become extremely marginal after ∼512 tokens. Even with a codebook size of 32, Phaedra already exhibits considerable gains compared to the FSQ baseline.
Dataset
Model
nMAE ↓
nRMSE ↓
r L1↓
r L2↓
Sentinel-2 L2A
Continuous
21.583 ± 60.032
30.169 ± 69.422
3.687 ± 6.950
5.790 ± 8.281
FSQ
31.650 ± 79.865
52.845 ± 113.701
5.880 ± 9.838
10.113 ± 15.647
Phaedra 4
23.719 ± 58.275
31.756 ± 67.894
4.572 ± 7.132
6.446 ± 8.490
Phaedra 8
30.178 ± 56.766
41.157 ± 65.523
6.394 ± 6.389
9.276 ± 7.392
Sentinel-2 L1C
Continuous
21.198 ± 50.801
33.021 ± 60.200
4.214 ± 6.814
7.303 ± 8.156
FSQ
22.098 ± 21.393
51.321 ± 47.126
5.904 ± 4.410
13.675 ± 9.907
Appendix
Table 23 : Comprehensive evaluation across Earth Observation and Satellite datasets. Metrics are reported as mean±std . Relative metrics (r L1 /r L2 ) for DEM and NDVI are omitted due to numerical instability caused by zero-valued pixels in the ground truth. Results in this table for Sentinel-2 L1C differ from those in the main text because these are for a large-scale global evaluation, while the main text presented results for specific regions.
nMAE ↓
nMSE ↓
r L1↓
r L2↓
Continuous
0.023 ± 0.000
0.034 ± 0.001
3.575 ± 0.144
5.254 ± 0.216
FSQ
0.062 ± 0.001
0.091 ± 0.001
9.567 ± 0.330
14.003 ± 0.482
Phaedra 4
0.042 ± 0.000
0.059 ± 0.001
6.391 ± 0.250
9.117 ± 0.362
Phaedra 8
0.156 ± 0.003
0.231 ± 0.006
23.912 ± 1.098
35.491 ± 1.679
Appendix
Table 24 : ERA5 reanalyses: Zonal and meridional winds. Metrics measured in mean ± std.
nMAE ↓
nMSE ↓
r L1↓
r L2↓
Continuous
0.022 ± 0.001
0.037 ± 0.001
0.080 ± 0.002
0.133 ± 0.004
FSQ
0.062 ± 0.002
0.095 ± 0.002
0.224 ± 0.006
0.342 ± 0.007
Phaedra 4
0.038 ± 0.001
0.063 ± 0.002
0.138 ± 0.004
0.226 ± 0.007
Phaedra 8
0.121 ± 0.004
0.192 ± 0.006
0.437 ± 0.015
0.690 ± 0.022
Appendix
Table 25 : ERA5 reanalyses: Temperature. Metrics measured in mean ± std.
Model
Density
Velocity (avg.)
Pressure
FSQ
0.98
5.94
0.79
Phaedra
0.52
2.81
0.37
Appendix
Table 26 : Relative L1 error for the 3D CEU RC reconstructions.
Figure 13 : CEU RC, ρ : Ground truth and reconstructions for density at the final timestep in the first trajectory of the CEU RC dataset.
Figure 14 : CEU KH, ρ : Ground truth and reconstructions for density at the final timestep in the first trajectory of the CEU KH dataset.
Figure 15 : CEU KH, p : Ground truth and reconstructions for pressure at the final timestep in the first trajectory of the CEU KH dataset.
Figure 16 : CEU AIR, ρ : Ground truth and reconstructions for density in the first sample of the CEU Airfoil dataset.
Figure 17 : AWA : Ground truth and reconstructions for the solution at the final timestep in the first trajectory of the Acoustic Wave dataset.
Figure 18 : CEU RC 512 , ρ : Ground truth and reconstructions for density at the final timestep in the first trajectory of the high-resolution CEU RC dataset.
Figure 19 : CEU KH 512 , ρ : Ground truth and reconstructions for density at the final timestep in the first trajectory of the high-resolution CEU KH dataset.
Figure 20 : CEU KH 512 , E : Ground truth and reconstructions for energy at the final timestep in the first trajectory of the high-resolution CEU KH dataset.
Figure 21 : Sentinel-2 L1C-Band 3 : Ground truth and reconstructions for the first sample of Band 3. All models are applied without any fine-tuning on Sentinel-2 or other earth observation data. We use a dataset-wide 0:1 normalization across all models.
Figure 22 : Global positions of the 10 locations reconstructed in Fig. 23 .
Figure 23 : Original (left) vs Reconstruction (right) for the Sentinel-2 RGB subset. Locations 1-10 are shown in order left-to-right, top-to-bottom.
Figure 28 : Examples of image reconstructions by Phaedra and Cosmos. Phaedra exhibits smoothing, especially visible under the 162 downsampling. This is primarily a result of the training process, as natural images have a frequency spectra which decays much slower than many PDEs. Alternatively, Cosmos 16 reconstructs sharper images, but exhibits failure modes such as misplacing sharp transitions (as visible in the chains and feathers) or completely removing details (e.g. the eye of the quail).
Model
Dataset
ρ
u
v
p
Average
Strategy
CI 95
FNO
KH
6.24
11.18
32.23
0.54
12.55
6-6-2
±0.36
FNO
RC
26.31
53.62
53.42
9.16
35.63
6-6-2
±0.69
FNO
RKH
10.08
26.36
25.67
4.54
16.66
6-6-2
±0.91
CNO
KH
5.06
9.10
25.93
0.51
10.15
2-step
±0.32
CNO
RC
22.65
44.44
44.66
7.94
29.92
2-step
±0.74
CNO
RKH
6.79
18.83
18.50
3.17
11.82
6-6-2
±0.82
Appendix
Table 27 : Full per-variable operator learning relative L1 errors at final time step ( t=0.7 ). Three rollout time-stepping strategies with short (2-step), medium (6-6-2 step), and long (direct) steps are tested, and the best results for each model are reported.
Model
Dataset
t=0.1
t=0.2
t=0.3
t=0.4
t=0.5
t=0.6
t=0.7
FNO
KH
6.63
7.43
8.23
9.20
10.34
11.60
13.10
CNO
KH
5.15
5.97
6.73
7.70
8.80
10.05
11.72
ViT
KH
5.20
5.53
6.06
6.81
7.67
8.53
9.71
Continuous Transformer
KH
6.15
6.08
6.32
6.86
7.59
8.40
9.42
VQ-VAE-2
KH
12.19
11.11
10.80
11.54
12.48
13.76
15.28
FSQ
KH
6.78
7.64
8.24
8.98
10.09
11.09
11.99
Appendix
Table 28 : Operator learning relative L1 errors under the direct-step strategy (no rollout) across timesteps.
Model Size
t=0.1
t=0.2
t=0.3
t=0.4
t=0.5
t=0.6
t=0.7
5M
16.94±3.61
19.59±4.86
24.19±5.99
28.68±6.74
30.88±7.10
33.51±9.22
36.16±9.38
38M
5.38±1.58
6.42±1.89
8.77±2.97
11.53±4.15
13.39±4.81
15.98±6.92
18.99±7.93
200M
3.72±1.23
4.34±1.45
6.03±2.38
8.10±3.38
9.88±4.06
12.10±6.02
14.89±7.45
Appendix
Table 29 : Performance across timesteps and model scales for the CEU RKH dataset.
Model
Dataset
Relative L1 Error
W1
FSQ
KH
5.84
1.71
FSQ
RC
29.16
21.71
FSQ
RKH
9.45
8.19
Phaedra
KH
4.75
1.21
Phaedra
RC
22.17
12.50
Phaedra
RKH
7.27
5.52
Appendix
Table 30 : Full masked autoencoding results. Average across all 4 physical channels. Inputs are 75% masked. Wasserstein-1 (W1) metrics are scaled by 103 .
Figure 29 : Visualizations of the discrete Phaedra tokens predicted by the transformer model across the three evaluation tasks. Each figure corresponds to the first test sample from the test set. All predictions are produced by a rollout from the initial conditions.
Figure 30 : Comparison of the predicted physical fields against the ground truth simulations for the CEU RC operator learning task for the first test sample. Each model is presented under its best strategy (d=direct, r2=2-step, 662=6-6-2 lead times; label = avg. rel L1 ).
Figure 31 : Comparison of the predicted physical fields against the ground truth simulations for the CEU RKH operator learning task for the first test sample. Each model is presented under its best strategy (d=direct, r2=2-step, 662=6-6-2 lead times; label = avg. rel L1 ).
Figure 32 : Masked Autoencoding (MAE) reconstructions for selected fields under 75% spatial masking. Each panel contrasts the masked input provided to the model with the predicted reconstruction and the unmasked ground truth.
Conventional patchified Transformers operate on uniform spatial partitions, distributing computational effort evenly across the domain irrespective of local features. This inflexible tokenization scheme is inherently limited in its ability to efficiently represent and process solutions to complex PDEs. To address this, we propose MeshTok, an adaptive mesh refinement (AMR)-inspired tokenization and sequence modeling framework. This method selectively refines spatial regions exhibiting sharp gradients, transient features, or multiscale structures, generating a heterogeneous set of multiscale tokens defined on a fixed simulation grid. These tokens are processed within a unified Transformer sequence, enabling the model to simultaneously capture coarse-grained global context and fine-grained local details without requiring specialized architectural components. Although adaptive refinement moderately increases token count, it promotes a more targeted allocation of computational resources to physically informative regions, which we view as a practical inductive bias rather than a formal optimality guarantee. Experimental evaluations across multiple PDE families and benchmark datasets demonstrate that MeshTok consistently improves the efficiency-accuracy trade-off compared to uniform-grid baselines. This suggests adaptive multiscale tokenization as a scalable and generalizable design principle for neural PDE modeling. Code is available at https://github.com/SCAILab-USTC/MeshTok.
Yanshun Zhao, Xiaoyu Peng, Jiamin Jiang +2
School of Mathematical Sciences, University of Science and Technology of China, Hefei 230026, China · Suzhou Institute for Advanced Research, University of Science and Technology of China, Suzhou 215123, China.
Transformer architectures have attracted increasing attention for solving partial differential equations (PDEs), owing to their flexibility in handling irregular discretizations and their ability to capture long-range physical dependencies. However, unlike discrete language tokens or fixed-resolution image patches, observed physical fields are finite samples of underlying infinite-dimensional functions. Consequently, effectively applying Transformers to PDEs requires a tokenizer that respects the functional nature of physical fields and constructs physically expressive tokens from arbitrary discretizations.To this end, we propose \methodname{Physics Transformer}, a function-projection-based Transformer architecture for physical field prediction. Physics Transformer treats a physical field as a continuous function and partitions its discretization into locality-preserving spatial patches. Within each patch, it dynamically learns a set of adaptive local basis functions and projects the sampled field onto these bases to obtain compact physics tokens. The resulting tokens capture diverse latent physical states while preserving fine-scale spatial structures, enabling efficient global interaction through factorized attention across space and physical states. The projected representation further supports efficient decoding at arbitrary query locations. Extensive experiments on diverse benchmarks, ranging from two-dimensional PDE dynamics to industrial-scale three-dimensional CFD simulations, demonstrate that Physics Transformer accurately captures fine-grained physical structures and achieves state-of-the-art predictive performance. These results establish function projection as a practical and effective foundation for designing Transformer architectures for PDE solving.
Guoze Sun, Rui Zhang, Jiankai Tang +4
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · School of Mechanics and Engineering Science, Peking University, Beijing, China
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compressed to extremely low token budgets. In this paper, we present ZipTok3D, a 3D tokenizer designed for high-fidelity reconstruction from extremely short token sequences. Its key idea is to organize object geometry into progressively informative global-token prefixes and unfold these compact representations through iterative decoding. Specifically, nested dropout randomly truncates the latent sequence after encoding during training and requires each retained prefix to reconstruct the complete object, thereby prioritizing essential geometric information in the leading tokens. The decoder then repeatedly applies a parameter-shared Transformer block to recover fine-grained geometry from each prefix without a separate generative sampling stage. With the same token dimension, ZipTok3D achieves reconstruction quality comparable to the 32-token COD-VAE baseline using only one token on ShapeNet and four on TRELLIS, yielding 32× and 8× shorter token sequences, respectively.
Mingda Lin, Weijie Wang, Zeyu Zhang +7
1Zhejiang University · 2Monash University · University of Adelaide