Atomizer-IO: Beyond Pixels, Patches and Grids
Organizations: INRIA, Montpellier, France · LIPADE, Université Paris Cité, Paris, France · DM3L, University of Zurich, Zurich, Switzerland · CIRAD, Montpellier, France
Abstract
Most vision architectures assume that observations lie on a regular grid, an effective abstraction for natural images but a restrictive one for sensing data whose channels, temporal sampling, spatial resolution, and geometry can vary. Generic set-based architectures remove the grid, but also remove useful spatial inductive biases. We introduce Atomizer-IO, an architecture that places observations first and derives structure from their physical relationships. Building on top of an atomic representation of the data, each observation is described by its measurement and acquisition metadata, while local cross-attention maps observations to anchor points that can be arbitrarily placed. We evaluate this design by progressively relaxing the grid assumption, from varying input raster configurations and incomplete channel sets to flexible output density and, ultimately, inputs without a raster grid. Atomizer-IO is competitive with flexible EO-specific architectures on most tasks, while offering post-training control over inference cost and competitive compute--performance trade-offs. The same formulation extends without architectural redesign to unordered 3D point clouds, showing that the atomic interface generalizes beyond regular raster inputs. These results suggest that pixels, patches, and grids do not need to define the interface of a sensing architecture.
Figures & tables
| Dataset | Sensor | T | C | Geometry | GSD | Problem | Out. C |
| ForestNet ( Irvin et al., 2020 ) | Landsat-8 | 1 | 6 | 15 m | cls | 12 | |
| BurnScars ( Jakubik et al., 2023 ) | HLS | 1 | 6 | 30 m | seg | 2 | |
| EuroSAT ( Helber et al., 2019 ) | Sentinel-2 | 1 | 13 | 10 m | cls | 10 | |
| Cashew ( Jin et al., 2021 ) | Sentinel-2 | 1 | 13 | 10 m | seg | 7 | |
| Sen1Floods11 ( Rambour et al., 2020 ) | S1 + S2 | 1 | 15 | 10 m | seg | 2 | |
| xView2 ( Gupta et al., 2019 ) | RGB (VHR) | 2 | 3 | 0.5 m | seg | 5 |
| Model | EuroSAT | ForestNet | BurnScars | Sen1Floods | PASTIS | xView2 | Cashew | BioMassters |
| ResNet50 | 90.9 | 43.8 | 82.8 | 87.8 | 30.0 | 53.8 | 69.9 | 46.3 |
| ViT | 90.7 | 39.1 | 87.5 | 85.0 | 35.1 | 52.4 | 47.5 | 52.7 |
| RAMEN | 92.1 | 40.4 | 88.2 | 89.1 | 45.5 | 53.3 | 64.2 | 44.7 |
| UniverSat | 90.8 | 44.3 | 88.1 | 86.8 | 42.9 | 56.3 | 62.4 | 46.0 |
| Perceiver-IO | 90.9 | 35.4 | 83.4 | 92.2 | 17.9 | 43.8 | 26.2 | 52.2 |
| Atomizer-IO | 90.2 | 35.8 | 88.8 | 93.2 | 44.1 | 54.6 | 72.0 | 42.4 |
| Sen1Floods11 | BioMassters | |||||||||||
| Test Configuration | Atom. | Res. | ViT | Perc. | RAM. | Uni. | Atom. | Res. | ViT | Perc. | RAM. | Uni. |
| All bands | 92.9 | 88.1 | 84.5 | 92.1 | 87.0 | 88.7 | 41.5 | 47.6 | 52.9 | 52.0 | 43.4 | 45.7 |
| S2 only | 92.8 | 86.4 | 84.4 | 92.2 | 87.8 | 88.6 | 45.8 | 55.9 | 55.5 | 57.7 | 48.4 | 49.1 |
| S1 only | 80.0 | 66.4 | 75.3 | 77.2 | 75.1 | 79.4 | 48.1 | 52.1 | 55.5 | 58.6 | 49.2 | 49.1 |
| RGB only | 44.8 | 43.7 | 43.8 | 43.8 | 44.5 | 43.8 | 71.0 | 73.8 | 71.4 | 70.4 | 71.9 | 70.8 |
| No SWIR | 87.8 | 49.0 | 75.1 | 81.8 | 63.7 | 85.2 | 43.0 | 49.1 | 54.0 | 53.4 | 46.0 | 46.7 |
| DALES | |
| Method | mIoU |
| KPConv ( Thomas et al., 2019 ) | 81.1 |
| SPT ( Robert et al., 2023 ) | 79.6 |
| PointNet++ ( Qi et al., 2017b ) | 68.3 |
| ConvPoint ( Boulch, 2020 ) | 67.4 |
| Atomizer-IO | 66.7 |
| MNIST | Sen1Fl11 | |
| No RoPE | 29.48 | 92.42 |
| RoPE | 99.37 | 93.17 |
| Background kept (%) | Atomizer-IO | ViT |
|---|---|---|
| 99.26 | 98.11 | |
| 98.72 | 97.65 | |
| 98.37 | 94.85 | |
| 94.93 | 72.17 | |
| 74.36 | 37.24 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| MNIST | Cashew | PASTIS | BurnScars | Sen1Floods11 | |
| No RoPE | 29.48 | 62.65 | 40.65 | 86.16 | 92.42 |
| RoPE | 99.37 | 72.01 | 44.10 | 88.81 | 93.17 |
| Sen1Floods11 | Cashew | PASTIS | BurnScars | |
| Without local observations | 89.26 | 70.58 | 43.30 | 88.20 |
| Full model | 93.17 | 72.01 | 44.10 | 88.81 |
| Interaction bound | Equiv. global drop | Context/query | |
| Global | 0% | ||
| Local ( ) | 99.987% | ||
| Local ( ) | 99.975% | ||
| Local ( ) | 99.949% |
| Full decode calls | Best case | Worst case | |
| Dense | |||
| Zone-probe | |||
| Quadtree |
| Architecture | Classification | Segmentation | Multitemporal seg. |
| ResNet50 | 23.6 | 37.3 | 37.3 |
| ViT-Small | 21.7 | 30.3 | 31.5 |
| Perceiver IO | 19.7 | 34.5 | 34.5 |
| RAMEN | 23.1 | 33.7 | 33.7 |
| UniverSat | 36.1 | 36.1 | 36.1 |
| Atomizer-IO | 19.2 | 33.5 | 34.8 |
| RAMEN | UniverSat | |||
| Dataset | Enc. res. (m) | Enc. res. (m) | Subpatch (px) | Dec. stride (px) |
| ForestNet | 40 | 40 | 1 | – |
| BurnScars | 120 | 120 | 1 | 4 |
| EuroSAT | 40 | 40 | 1 | – |
| Cashew | 40 | 40 | 1 | 1 |
| Sen1Floods11 | 40 | 40 | 1 | 4 |
| Dataset | Batch size |
| ForestNet | 8 |
| BurnScars | 8 |
| EuroSAT | 8 |
| Cashew | 8 |
| Sen1Floods11 | 4 |
| xView2 | 8 |
| Hyperparameter | 2D | 3D |
| Latent dimension | 512 | 768 |
| Global latents | 128 | 128 |
| Attention heads | 8 | 8 |
| Decoder neighbors | 9 | 9 |
| Latent layout | Hexagonal | Hexagonal |
| Other 2D: PASTIS: |
| Feature | Frequency bands | Maximum frequency / period |
| Position | 16 | 16 |
| Resolution | 2 | 2 |
| Acquisition time | 12 | 365 days |
| Measurement value | 8 | 8 |
| Parameter | Atomizer-IO | ViT |
| Parameters (total) | 7.40M | 7.44M |
| Hidden / latent dimension | 256 | 224 |
| Encoder / processor depth | 1 cross-attn + 4 self-attn | 12 transformer blocks |
| Attention heads | 4 (cross) / 8 (self) | 4 |
| MLP dimension | 768 | 896 ( ) |
| Positional encoding | Relative (RoPE, self-attn) | Absolute (learned) |
| Method | Mean | Ground | Buildings | Cars | Trucks | Poles | Power lines | Fences | Vegetation |
| KPConv | 0.811 | 0.971 | 0.966 | 0.853 | 0.419 | 0.750 | 0.955 | 0.635 | 0.941 |
| PointNet++ | 0.683 | 0.941 | 0.891 | 0.754 | 0.303 | 0.400 | 0.799 | 0.462 | 0.912 |
| ConvPoint | 0.674 | 0.969 | 0.963 | 0.755 | 0.217 | 0.403 | 0.867 | 0.296 | 0.919 |
| Atomizer-IO | 0.667 | 0.962 | 0.941 | 0.670 | 0.095 | 0.549 | 0.829 | 0.395 | 0.895 |
| SuperPoint | 0.606 | 0.947 | 0.934 | 0.629 | 0.187 | 0.285 | 0.652 | 0.336 | 0.879 |
| PointCNN | 0.584 | 0.975 | 0.957 | 0.406 | 0.048 | 0.576 | 0.267 | 0.526 | 0.917 |
| Stage | Symbol | Meaning |
| Input | Number of input atoms | |
| Measured value of atom | ||
| Number of metadata fields of an atom | ||
| Metadata fields of atom | ||
| Encoder of the measured value | ||
| Encoders of the metadata fields |