Localize Any Object in X-Ray Security Scans without Human Annotation
Organizations: University of Trento · Fondazione Bruno Kessler · Dalian University of Technology · Hefei University of Technology
Abstract
Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder the direct zero-shot transfer of dense perception foundation models trained on web-scale RGB data. Moreover, annotated X-ray data is scarce and requires expert labeling, limiting both the training of generalizable X-ray native models and the adaptation of RGB foundation models for X-ray data via fine-tuning. Given these challenges, the bright promise of highly generalizable perception models, enabled by data scaling laws in the RGB domain, remains largely out of reach for X-ray inspection. To this end, we introduce LAO-X, a self-supervised adaptation framework that Locates Any Object in X-ray scans using diverse synthesized image--annotation pairs with granularity-aware supervision. LAO-X first designs a saliency-guided X-ray object mining module to separate diverse object instances, which are then used for physics-guided synthesis in the absorbance domain. LAO-X further incorporates an occlusion-controlled curriculum strategy to fine-tune a Segment Anything Model 2 (SAM2) localizer, progressively adapting it to X-ray scans with increasing object counts and overlap levels. Experiments on six X-ray benchmarks show that LAO-X substantially improves category-agnostic localization, achieving 2% to 23% mAP gains over SAM2 and X-ray specific baselines in heavily cluttered scenarios, entirely without human-annotated labels.
Figures & tables
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Evaluation set | |
|---|---|
| DET-COMPASS, full set | 63.39 |
| DET-COMPASS, invisible subset | 26.69 |
| Method | |||||
|---|---|---|---|---|---|
| SAM2 | 0.03 | 0.02 | 0.02 | 0.34 | 0.36 |
| SAM3 | 4.18 | 2.13 | 2.27 | 7.74 | 31.35 |
| UnSAMv2 | 3.84 | 1.18 | 1.61 | 8.05 | 19.59 |
| LAO-X | 5.87 | 2.60 | 2.98 | 11.92 | 14.60 |