Authors: Hugo Riffaud de Turckheim, Sylvain Lobry, Nicolas Houdré, Damien Robert, Roberto Interdonato, Diego Marcos
Organizations: INRIA, Montpellier, France · LIPADE, Université Paris Cité, Paris, France · DM3L, University of Zurich, Zurich, Switzerland · CIRAD, Montpellier, France
Most vision architectures assume that observations lie on a regular grid, an effective abstraction for natural images but a restrictive one for sensing data whose channels, temporal sampling, spatial resolution, and geometry can vary. Generic set-based architectures remove the grid, but also remove useful spatial inductive biases. We introduce Atomizer-IO, an architecture that places observations first and derives structure from their physical relationships. Building on top of an atomic representation of the data, each observation is described by its measurement and acquisition metadata, while local cross-attention maps observations to anchor points that can be arbitrarily placed. We evaluate this design by progressively relaxing the grid assumption, from varying input raster configurations and incomplete channel sets to flexible output density and, ultimately, inputs without a raster grid. Atomizer-IO is competitive with flexible EO-specific architectures on most tasks, while offering post-training control over inference cost and competitive compute--performance trade-offs. The same formulation extends without architectural redesign to unordered 3D point clouds, showing that the atomic interface generalizes beyond regular raster inputs. These results suggest that pixels, patches, and grids do not need to define the interface of a sensing architecture.
Figures & tables
Figure 1: Sensing configurations considered in this work. All are processed through the same atomic interface, without rasterization, resampling, or modality-specific branches.
Figure 2: The Atomizer-IO architecture. Atoms are locally aggregated into a compact spatial representation, processed globally in latent space, and decoded through spatial output queries. The same local context-to-query operation is used for encoding and decoding, with spatial relationships defined in physical coordinates rather than by a fixed input grid.
Figure 3: Token construction in Atomizer-IO. 2D imagery and LiDAR share the same atomic interface of measurements and acquisition metadata; spatial latents, global latents, and output queries use learned initializations.
Figure 4: Spatial sampler. From left to right: query and context sets; construction of a local receptive field from spatial proximity, illustrated by the encoder’s nearest-latent assignment and resulting Voronoi partition; query-centered relative coordinates; local cross-attention.
Figure 5: Example spatial latent partition on DALES (Section 4.4 ). Colors show the local receptive fields induced by assigning each input atom to its nearest spatial latent.
Dataset
Sensor
T
C
Geometry
GSD
Problem
Out. C
ForestNet ( Irvin et al., 2020 )
Landsat-8
1
6
3322
15 m
cls
12
BurnScars ( Jakubik et al., 2023 )
HLS
1
6
5122
30 m
seg
2
EuroSAT ( Helber et al., 2019 )
Sentinel-2
1
13
642
10 m
cls
10
Cashew ( Jin et al., 2021 )
Sentinel-2
1
13
2562
10 m
seg
7
Sen1Floods11 ( Rambour et al., 2020 )
S1 + S2
1
15
5122
10 m
seg
2
xView2 ( Gupta et al., 2019 )
RGB (VHR)
2
3
5122
0.5 m
seg
5
Table 1: Datasets used for evaluation, spanning raster Earth observation and irregular LiDAR inputs. T : timesteps; C : input channels; Geometry : spatial structure of the input; GSD : ground sampling distance, undefined for point clouds; Out. C : output classes or regression targets.
Model
EuroSAT
ForestNet
BurnScars
Sen1Floods
PASTIS
xView2
Cashew
BioMassters ↓
ResNet50
90.9
43.8
82.8
87.8
30.0
53.8
69.9
46.3
ViT
90.7
39.1
87.5
85.0
35.1
52.4
47.5
52.7
RAMEN
92.1
40.4
88.2
89.1
45.5
53.3
64.2
44.7
UniverSat
90.8
44.3
88.1
86.8
42.9
56.3
62.4
46.0
Perceiver-IO
90.9
35.4
83.4
92.2
17.9
43.8
26.2
52.2
Atomizer-IO
90.2
35.8
88.8
93.2
44.1
54.6
72.0
42.4
Table 2: Single-task performance. Best in bold , second best underlined . Classification uses macro F1, segmentation mIoU, and BioMassters RMSE (Mg/ha, ↓ ).
Sen1Floods11 ↑
BioMassters ↓
Test Configuration
Atom.
Res.
ViT
Perc.
RAM.
Uni.
Atom.
Res.
ViT
Perc.
RAM.
Uni.
All bands
92.9
88.1
84.5
92.1
87.0
88.7
41.5
47.6
52.9
52.0
43.4
45.7
S2 only
92.8
86.4
84.4
92.2
87.8
88.6
45.8
55.9
55.5
57.7
48.4
49.1
S1 only
80.0
66.4
75.3
77.2
75.1
79.4
48.1
52.1
55.5
58.6
49.2
49.1
RGB only
44.8
43.7
43.8
43.8
44.5
43.8
71.0
73.8
71.4
70.4
71.9
70.8
No SWIR
87.8
49.0
75.1
81.8
63.7
85.2
43.0
49.1
54.0
53.4
46.0
46.7
Table 3: Performance under flexible input composition on Sen1Floods11 (mIoU, %) and BioMassters (RMSE, Mg/ha). All models are trained under the same band-dropout protocol.
Figure 6: Compute cost (GFLOPs) versus segmentation performance (mIoU). Dense, Zone-probe, and Quadtree decoding vary inference cost through different output query strategies, using the same Atomizer-IO checkpoint without retraining.
DALES
Method
mIoU
KPConv ( Thomas et al., 2019 )
81.1
SPT ( Robert et al., 2023 )
79.6
PointNet++ ( Qi et al., 2017b )
68.3
ConvPoint ( Boulch, 2020 )
67.4
Atomizer-IO
66.7
Table 4: DALES and FRACTAL segmentation results (mIoU, %).
MNIST
Sen1Fl11
No RoPE
29.48
92.42
RoPE
99.37
93.17
Table 5: RoPE ablation. Accuracy on MNIST and mIoU on Sen1Floods11.
Background kept (%)
Atomizer-IO
ViT
100
99.26
98.11
75
98.72
97.65
50
98.37
94.85
25
94.93
72.17
10
74.36
37.24
Table 6: MNIST accuracy (%) under unseen input geometries. Digit pixels are always retained; both models use the same checkpoint trained on complete inputs.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Detailed Atomizer-IO architecture. The encoder aggregates local neighborhoods of input atoms into spatial latents through the spatial sampler; the processor applies latent self-attention with optional temporal aggregation; and the decoder maps local latent neighborhoods to spatial output queries, optionally conditioning first on raw observations at the target location. Relative positional information is introduced locally, and MLP and attention weights are shared across spatial neighborhoods.
Figure 8: Temporal aggregation ablation on PASTIS as the number of input timesteps increases.
MNIST
Cashew
PASTIS
BurnScars
Sen1Floods11
No RoPE
29.48
62.65
40.65
86.16
92.42
RoPE
99.37
72.01
44.10
88.81
93.17
Appendix
Table 7: Effect of spatial RoPE across tasks. MNIST reports accuracy (%), while the EO datasets report mIoU (%).
Figure 9: Relative degradation when spatial RoPE is removed, computed as 100(RoPE−NoRoPE)/RoPE .
Sen1Floods11
Cashew
PASTIS
BurnScars
Without local observations
89.26
70.58
43.30
88.20
Full model
93.17
72.01
44.10
88.81
Appendix
Table 8: Effect of conditioning output queries on the raw observations at their target location. All values are mIoU (%).
Figure 10: Relative degradation when local observation conditioning is removed, computed as 100(Full−NoLocal)/Full .
Interaction bound
Equiv. global drop
Context/query
Global
LN≈7.86×109
0%
N
Local ( m=500 )
≤Lm=1.0×106
99.987%
≤500
Local ( m=1000 )
≤Lm=2.0×106
99.975%
≤1000
Local ( m=2000 )
≤min(N,Lm)≈3.93×106
99.949%
≤2000
Appendix
Table 9: Cross-attention complexity for global and local spatial sampling on Sen1Floods11 ( N≈3.9×106 atoms, L=2000 spatial latents). The equivalent global drop rate is 1−m/N , the fraction of atoms that global cross-attention would need to discard to match a budget of m context elements per latent.
Full decode calls
Best case
Worst case
Dense
M
M
M
Zone-probe
kpL+∑ℓ∈hard∣Wℓ∣
O(kpL)
M
Quadtree
≲kp×(levels×active cells)
O(sstart2kpHW)
M
Appendix
Table 10: Decode cost of the three inference strategies. M is the number of output queries, L the number of spatial latents, kp the number of probes per region, and sstart the initial quadtree cell size. Both adaptive strategies reduce to dense decoding in the worst case when no region can be treated as homogeneous.
Figure 11: Sensitivity of segmentation performance and inference cost to three architectural parameters: atoms per latent, encoder cross-attention budget, and decoder neighborhood size. Results are shown on Sen1Floods11 and BurnScars for Dense, Zone-probe, and Quadtree decoding.
Figure 12: LiDAR return encoding. The pair (a,b) describes the relative position of each observation within its pulse’s return sequence. Normalization by the per-pulse return count R makes the representation independent of the sensor-specific maximum number of returns while preserving the original (r,R) attributes.
Architecture
Classification
Segmentation
Multitemporal seg.
ResNet50
23.6
37.3
37.3
ViT-Small
21.7
30.3
31.5
Perceiver IO
19.7
34.5
34.5
RAMEN
23.1
33.7
33.7
UniverSat
36.1
36.1
36.1
Atomizer-IO
19.2
33.5
34.8
Appendix
Table 11: Number of trainable parameters for the architectures used in the 2D experiments. Values are given in millions of parameters. ViT refers to ViT-Small.
RAMEN
UniverSat
Dataset
Enc. res. (m)
Enc. res. (m)
Subpatch (px)
Dec. stride (px)
ForestNet
40
40
1
–
BurnScars
120
120
1
4
EuroSAT
40
40
1
–
Cashew
40
40
1
1
Sen1Floods11
40
40
1
4
Appendix
Table 12: Architecture-specific spatial hyperparameters used for RAMEN and UniverSat. Encoding resolutions are expressed in physical units.
Dataset
Batch size
ForestNet
8
BurnScars
8
EuroSAT
8
Cashew
8
Sen1Floods11
4
xView2
8
Appendix
Table 13: Batch size used for each dataset. The same batch size is used across architectures within a dataset.
Hyperparameter
2D
3D
Latent dimension
512
768
Global latents
128
128
Attention heads
8
8
Decoder neighbors
9
9
Latent layout
Hexagonal
Hexagonal
(k,m)
Other 2D: {(1000,1000),(1500,1500),(2000,2000)} PASTIS: {(100,100),(250,250),(350,350)}
(1000,1000)
Appendix
Table 14: Main Atomizer-IO architectural hyperparameters used for the 2D and 3D experiments. Here, k denotes the target number of atoms per spatial latent and m the maximum number of atoms sampled per latent for local cross-attention.
Feature
Frequency bands
Maximum frequency / period
Position
16
16
Resolution
2
2
Acquisition time
12
365 days
Measurement value
8
8
Appendix
Table 15: Fourier feature parameters used by Atomizer-IO and Perceiver IO.
Figure 13: MNIST sparsification protocol. Models are trained on complete images. At evaluation, pixels with value v≤0.5 are progressively removed, while pixels with value v>0.5 are always retained.
Parameter
Atomizer-IO
ViT
Parameters (total)
7.40M
7.44M
Hidden / latent dimension
256
224
Encoder / processor depth
1 cross-attn + 4 self-attn
12 transformer blocks
Attention heads
4 (cross) / 8 (self)
4
MLP dimension
768
896 ( 4× )
Positional encoding
Relative (RoPE, self-attn)
Absolute (learned)
Appendix
Table 16: Architecture and training hyperparameters for the capacity-matched Atomizer-IO and ViT models used in the MNIST sparsification experiment.
Method
Mean
Ground
Buildings
Cars
Trucks
Poles
Power lines
Fences
Vegetation
KPConv
0.811
0.971
0.966
0.853
0.419
0.750
0.955
0.635
0.941
PointNet++
0.683
0.941
0.891
0.754
0.303
0.400
0.799
0.462
0.912
ConvPoint
0.674
0.969
0.963
0.755
0.217
0.403
0.867
0.296
0.919
Atomizer-IO
0.667
0.962
0.941
0.670
0.095
0.549
0.829
0.395
0.895
SuperPoint
0.606
0.947
0.934
0.629
0.187
0.285
0.652
0.336
0.879
PointCNN
0.584
0.975
0.957
0.406
0.048
0.576
0.267
0.526
0.917
Appendix
Table 17: Per-class IoU on DALES. Values are reported as fractions. Baseline values are those reported for the corresponding methods on DALES.
Stage
Symbol
Meaning
Input
N
Number of input atoms
vi
Measured value of atom i
J
Number of metadata fields of an atom
ui1,…,uiJ
Metadata fields of atom i
ϕ0(⋅)
Encoder of the measured value
ϕ1(⋅),…,ϕJ(⋅)
Encoders of the metadata fields
Appendix
Table 18: Notation used throughout the methodology.
Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where input modalities, scales, and resolutions vary widely. We introduce UniverSat, a ViT-style backbone built around a Universal Patch Encoder that maps patches from arbitrary spatial, spectral, and temporal resolutions, and from both optical and non-optical sensors, into a shared embedding space with a shared set of weights. This enables training a single model on heterogeneous multimodal corpora via self-supervision, yielding robust, sensor-agnostic spatial features. We validate this approach with strong results across classification and segmentation on standard EO benchmarks from GeoBench, PANGEABench, and SpectralEarth. Our code and models are available at https://github.com/gastruc/UniverSat.
Yohann Perron, Guillaume Astruc, Nicolas Gonthier +2
LIGM, Ecole Nationale des Ponts et Chaussées, IP Paris, Univ Gustave Eiffel, CNRS · EFEO · LASTIG, Univ Gustave Eiffel, IGN, ENSG +2
When humans see a bird, they recognize far more than just "bird" -- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every bird they have ever seen. We ask whether a self-supervised visual model can discover the same compositional structure on its own. To this end, we propose RATS (Register Attention Transformers), which decomposes the classification token into N learnable register tokens that route patch information through an L->N->N->L bottleneck via a three-step compress-communicate-broadcast attention. The N registers are partitioned across the H attention heads, so that registers assigned to different heads do not interact with each other. Without auxiliary losses or part annotations, each register spontaneously specializes into a proto-semantic region whose emerging structure resembles object parts. RATS surpasses all baselines by +12 mIoU on average across five segmentation benchmarks, with consistent gains on ADE20K (+1.11 mIoU) and COCO (+0.2 AP^m). Its register dictionary further exhibits part-level consistency and semantic proximity across related categories. Our results suggest that RATS may provide a useful architectural prior for structured and interpretable visual representation learning.
Timing Yang, Predrag Neskovic, Jansen Seheult +4
1Johns Hopkins University · 2Office of Naval Research, Arlington, VA · Department of Laboratory Medicine and Pathology, Mayo Clinic,FLOPs (G)MN, USA
We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across diverse domains, spanning remote sensing, outdoor LiDAR, indoor RGB-D sequences, object-centric CAD models, and point clouds lifted from RGB-only videos. Despite their distinct sensing geometries, densities, and priors, Utonia learns a consistent representation space that transfers across domains. This unification improves perception capability while revealing intriguing emergent behaviors that arise only when domains are trained jointly. Beyond perception, we observe that Utonia representations can also benefit embodied and multimodal reasoning: conditioning vision-language-action policies on Utonia features improves robotic manipulation, and integrating them into vision-language models yields gains on spatial reasoning. We hope Utonia can serve as a step toward foundation models for sparse 3D data, and support downstream applications in AR/VR, robotics, and autonomous driving.
Yujia Zhang, Xiaoyang Wu, Yunhan Yang +6
The University of Hong Kong · Xiaomi · The Chinese University of Hong Kong