A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-K proposal coverage Cτ: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top K regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. Cτ falls while in-domain accuracy rises, on four architectures and three domains, by 5.12 to 63.35 points on boxes above 1024 px2. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs 87% of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most 2.47 points of in-domain accuracy. Seeing it costs one extra evaluation pass.
Figures & tables
Figure 1: Fine-tune a detector on a narrow set of classes and it gets worse at finding everything else. The test that signs it off cannot show you. 1 OWLv2, five seeds, on the 21,861 COCO objects no deployment vocabulary names. One line is the share of them it still puts a box on; the other is its detection AP on the same boxes. Both run 0 to 100 but are not the same unit, so the two percentages give what each has lost of its own starting value (§ 4.1 ). Below, the same two checkpoints on one image. IoU measures how much a predicted box overlaps the true one, so the photographs show the boxing , the half that largely survives. The pair is the median of 586 qualifying boxes, not an extreme one (§ 4.3 ). 2 The in-domain test set holds no such object, so it cannot report any of this. 3 What we count, at every budget. The gap is still open at 3000 boxes, so the budget of 300 used elsewhere does not create it (§ ). Panels 1 and 2 are the problem and panel 3 the instrument; the repair is Figure 3 .
claim
replicated over
where
What happens
∙ A detector fine-tuned on a narrow vocabulary puts a box on fewer objects outside it, while its in-domain accuracy improves
∙ In-domain accuracy does not determine held-out coverage: across five seeds of one recipe it varies by 1.54 mAP while coverage varies by 2.25 points, and the most accurate seed is the least covered
5 seeds within one recipe; a constructed cross-recipe pair differs by 11.37 points, and a disjoint-support argument covers any in-domain functional
§ 3.2
★ The name goes before the box. After adaptation the model still puts a region on about three quarters of these objects and keeps an eighth of its AP on them, and already at a shallower freeze depth every seed has given up more of its naming than of its boxing
1 arch, the only one scorable by name; 1 domain; all 6 rungs of its ladder, 3 or 5 seeds each but one, every run agreeing at both thresholds and every replicated rung clear of zero
§ 4.1
Which objects it happens to
★ What is damaged is what the checkpoint already had. Categories the pretrained model never learned do not lose coverage; they gain it
3 arch, 752 LVIS categories, 5 seeds; all six cells clear of zero, every seed sharing its cell’s sign
§ 5.3
Table 1: Everything this paper claims, and what each claim was tested on. Ten claims in the four groups the contributions are stated in; ★ marks the five that are new here and ∙ the rest, which confirm or adapt a result the literature already carries (§ 2 ). Every row’s principal measurement rests on five or more independent seeds ; where a row adds a supporting leg carrying fewer, the middle column or the section it names gives the count. The middle column says how widely each claim was tested. Four further claims, resting on less and argued only in the supplement, are in Table .
term
what it denotes
held-out top- K proposal coverage
What we measure. Of the objects whose class the deployment vocabulary never names, the share that at least one of the detector’s top K boxes still lands on (Equation 2 ). Held out means held out of the deployment vocabulary , not of pretraining: the pretrained model was trained on every one of these classes, so this is what a checkpoint keeps , not what it generalises to
class-agnostic localisation
What it stands in for. Putting a box on an object whose name the model was never given. Not measured directly, and coverage can record a miss while this is intact
open-world proposal recall
Where it comes from. Average recall on categories held out of training , an established measurement ( Kim et al, 2022 ; Konan et al, 2022 ) . We keep it and change its use: from scoring a training recipe to reading one checkpoint before and after adaptation
unseen-vocabulary recognition
A different thing. Whether the model can still name categories withheld from adaptation. OWLv2 only, in mAP, never mixed into the coverage axis (§ )
Table 2: Four quantities this paper keeps apart , because they are easy to confuse and only the first is measured here. Table 3 sets it against the instruments that might have been used in its place.
needs
read before the
instrument
a name
class score
categories it scores
read from the detections a model emits
detection AP
yes
no
the ones it names
base/novel split AP ( Wu et al, 2025 )
yes
no
the ones it names
TIDE, Hoiem
yes
no
the ones it names
read from the regions it proposes
Table 3: Why the standard instruments cannot report this loss. Everything above the rule is read from the detections a model finally emits, so it can speak only about categories the model still has a name for. A narrow fine-tune leaves it 22 . Hoiem et al (2012) and Bolya et al (2020) come closest (§ 2 ), but both take their bins after ranking and non-maximum suppression, so a region the model still produces but no longer ranks into its output counts as missed there and as present here (§ 5.1 ). The last two rows read the same quantity of different categories : open-world proposal recall scores what a recipe never learned, Cτ what this checkpoint knew and was not asked about, its top K regions taken by the architecture’s own selection rule.
novel
control
novel,
architecture
condition
trainable
IoU 0.5
IoU 0.7
IoU 0.5
IoU 0.7
unfilt. 0.5
Faster R-CNN
pretrained (COCO)
—
95.10
84.26
94.11
84.00
88.54
tabletop-22, +RPN
14.6M
91.81 ± 0.23
77.32 ± 0.11
95.16 ± 0.11
84.30 ± 0.20
—
tabletop-22, all
41.2M
85.14 ± 1.00
66.58 ± 1.44
85.63 ± 1.11
71.59 ± 1.16
74.96
cardd-22, all
41.2M
78.71 ± 0.51
54.17 ± 0.99
75.83 † ± 1.80
57.82 † ± 1.64
62.78
OWLv2
pretrained
—
97.96
90.06
97.84
91.19
92.99
Table 4: The share of objects each detector still puts a box on, before and after fine-tuning , using its own top 300 regions and no class label. novel is the 21,861 COCO boxes no deployment vocabulary names, control the 2,225 that tabletop-22 does, so control is a naming contrast on the tabletop-22 rows only and † marks every cell where it is not. The last column drops the 1024px2 size floor (Table ). Each ± is one standard deviation over five seeds ; a cell without one is a single run, and --- means not measured, or under trainable that the row trains nothing. Seeds spread widest on OWLv2’s fine-tuned row, where the 95% interval is half the drop it reports. Grounding DINO: Table .
Figure 2: Adaptation damages what the checkpoint already had, and leaves alone what it never had. The 752 LVIS categories split by whether the pretrained model was trained to detect them. Both groups sit outside every deployment vocabulary and are read by the same scorer on the same images, so the split is on pretraining exposure alone. Bars are the mean over five fine-tuned seeds, whiskers a 95% interval; every bar clears zero and every seed shares its bar’s sign (Table ).
pair
domain
raw
partial
both
p
Faster R-CNN / YOLOX
tabletop-22
0.528
0.507
0.594
0.0001
cardd-22
0.626
0.581
0.599
0.0001
Faster R-CNN / RF-DETR
tabletop-22
0.482
0.323
0.324
0.0001
cardd-22
0.536
0.412
0.418
0.0001
YOLOX / RF-DETR
tabletop-22
0.319
0.180
0.191
0.0009
cardd-22
0.389
0.264
0.270
0.0001
Table 5: Different detectors lose the same categories. Each cell correlates two architectures’ per-category drops over the 333 LVIS categories with at least ten boxes, at IoU 0.5 . raw is that correlation; partial first removes each architecture’s own pretrained coverage, and both removes both. p comes from a 10,000 -permutation null that shuffles category labels within one architecture. At IoU 0.7 five of the six cells are larger, the exception being Faster R-CNN against YOLOX on cardd-22 (Table ).
architecture
domain
rungs read
shallowest
worst rung
full
worst is deepest
Faster R-CNN
tabletop-22
5 of 7
−6.94 ( +RPN )
−17.67 ( full )
−17.67
yes
cardd-22
5 of 7
−36.98 ( +RPN )
−36.98 ( +RPN )
−30.08
no
OWLv2
tabletop-22
6 of 6
−0.70 ( heads )
−52.24 ( full )
−52.24
yes
RF-DETR
tabletop-22
6 of 6
−2.01 ( heads )
−10.43 ( dec3 )
−5.12
no
cardd-22
6 of 6
−6.94 ( heads ) †
−43.60 ( dec1 ) †
−12.00
no
YOLOX
tabletop-22
5 of 5
−0.04 ( preds )
−12.87 ( full )
−12.87
yes
Table 6: Freezing more of the network does not reliably protect it : for three of these seven ladders the deepest stage was not the most damaging. Cells are the change in coverage at IoU 0.7 against that architecture’s own pretrained model, so rows are not comparable with one another. rungs read leaves out stages whose trained parameters the objectness score cannot see, which must measure −0.00 (§ ). † marks a rung below 10 in-domain mAP (§ ); those runs are kept here, and dropping them moves two rungs. Five seeds throughout except OWLv2 below full and Faster R-CNN’s interior rungs. Full ladders, and the two rungs: § .
Figure 3: What mixing the pretrained weights back in buys, and what it costs , Faster R-CNN. Each point is one mixing fraction α , running from the fine-tuned model at α=0 to the pretrained one at α=1 ; four fractions carry a marker of their own, keyed above, and the rest are open dots. Up is coverage recovered, left is in-domain accuracy given up, so a curve that rises steeply is buying cheaply. Bars are one standard deviation on each axis. The fractions are not measured on a common set of seeds : five carry α∈{0,0.25,0.5} , three carry {0.2,0.4,0.6} , and the two nearest the pretrained model are single evaluations, marked by a cross. The two domains pay very different prices: tabletop-22 turns almost vertically, cardd-22 trades the two roughly in proportion.
Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures. Prior metacognitive methods learn logical rules that flag a model's errors, but rely on hand-authored domain-knowledge cues (object-size priors, segmentation masks) that do not transfer to genuinely novel scenes. We show that this metacognitive layer can be learned without any domain knowledge by exploiting vector-space geometry: per-model Label Vector Pools (LVP), built from each model's own training embeddings, yield error-detection rules from the geometry of detections relative to training-determined prototypes, reaching parity with domain-knowledge rules to within 0.002 every F1 on test set. Because the approach remains neurosymbolic, these geometric rules share a single logical framework and can still be complemented by domain knowledge when available. We frame the fusion of multiple imperfect ViT-based detectors as a consistency-based abduction problem solved at test time by an exact Integer Program (IP) and a polynomial-time heuristic. On an aerial-imagery benchmark of 15 weather-shifted test sets and six ViT detectors, our domain-knowledge-free layer matches the strongest majority-vote variant on clean data (within 0.005 F1) and, unlike every majority-vote baseline, retains its performance under a coordinated label-flipping attack: at a 90% flip rate it averages 0.42 F1 versus 0.35 for MV-Plurality (a 22% relative gain) and attains the highest F1 on \emph{every} test set once the flip rate exceeds 0.4
Mario Leiva, Yue Ma, Qinru Qiu +2
1DCIC, Universidad Nacional del Sur (UNS) & ICIC (UNS-CONICET), Bahía Blanca, Argentina · 2Syracuse University, Syracuse, NY USA
Selecting a zero-shot out-of-distribution (OOD) detector for a new deployment is typically based on benchmark rankings, implicitly assuming that the highest-ranked detector will transfer across domains. We show that this assumption does not hold. Through a controlled portability audit across seventeen in-distribution datasets, three vision-language models, and seven representative zero-shot OOD detectors, we find that detector rankings reverse across deployments, every detector exceeds 80% FPR95 on at least one domain, and the preferred detector depends on both the in-distribution data and the underlying VLM. We trace these reversals to complementary evidence channels in vision-language logits. Corpus-free detectors rely on different combinations of absolute match level and relative or spatial sharpness, while WordNet-based methods additionally depend on external semantic coverage. A simple proposition shows that level and sharpness cannot generally be recovered from one another, explaining why no single detector transfers reliably across deployments. Motivated by this diagnosis, we introduce the Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles. Controls replacing these channels with entropy, logit variance, or random noise do not reproduce the gains. Without OOD samples, auxiliary corpora, or learned fusion, CEG reduces detector sensitivity and improves GL-MCM from 38.1 to 28.8 and MCM from 42.6 to 30.5 family-balanced FPR95.
Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Stephen Gould +1
University of Adelaide · 3Naval Group Pacific · 2Australian National University
Out-of-distribution (OOD) detection identifies test samples that fall outside a model's training distribution, a capability critical for safe deployment in high-stakes applications. Standard OOD detectors are trained on a specific in-distribution (ID) dataset and detect deviations from that single domain. In contrast, we study few-shot cross-domain OOD detection: given a \emph{single} pre-trained model, can we perform OOD detection on \emph{arbitrary} new ID-OOD task pairs using only a handful of ID samples at inference time, with no additional training? We propose \textbf{UFCOD}, a unified framework that achieves this goal through information-geometric analysis of diffusion trajectories. Our key insight is that diffusion noise predictions are score functions (gradients of log-density), and we extract two energy features: \emph{Path Energy} (integrated score magnitude) and \emph{Dynamics Energy} (score smoothness), that form a discrete Sobolev norm capturing how samples interact with the learned diffusion process. The central contribution is a \textbf{train-once, deploy-anywhere} paradigm: a diffusion model trained on a single dataset (e.g., CelebA) serves as a universal feature extractor for OOD detection across semantically unrelated domains (e.g., CIFAR-10, SVHN, Textures). At deployment, each new task requires only ∼100 unlabeled ID samples for inference: no retraining, no fine-tuning, no task-specific adaptation. Using 100 ID samples per task, UFCOD achieves 93.7% average AUROC across 12 cross-domain benchmarks, competitive with methods trained on 50k--163k samples, demonstrating ∼500× improvement in sample efficiency. See our code in https://github.com/lili0415/UFCOD.
Shawn Li, You Qin, Jiate Li +4
University of Southern California · National University of Singapore · Amazon