A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-K proposal coverage Cτ: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top K regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. Cτ falls while in-domain accuracy rises, on four architectures and three domains, by 5.12 to 63.35 points on boxes above 1024 px2. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs 87% of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most 2.47 points of in-domain accuracy. Seeing it costs one extra evaluation pass.
Figures & tables
Figure 1: Fine-tune a detector on a narrow set of classes and it gets worse at finding everything else. The test that signs it off cannot show you. 1 OWLv2, five seeds, on the 21,861 COCO objects no deployment vocabulary names. One line is the share of them it still puts a box on; the other is its detection AP on the same boxes. Both run 0 to 100 but are not the same unit, so the two percentages give what each has lost of its own starting value (§ 4.1 ). Below, the same two checkpoints on one image. IoU measures how much a predicted box overlaps the true one, so the photographs show the boxing , the half that largely survives. The pair is the median of 586 qualifying boxes, not an extreme one (§ 4.3 ). 2 The in-domain test set holds no such object, so it cannot report any of this. 3 What we count, at every budget. The gap is still open at 3000 boxes, so the budget of 300 used elsewhere does not create it (§ ). Panels 1 and 2 are the problem and panel 3 the instrument; the repair is Figure 3 .
claim
replicated over
where
What happens
∙ A detector fine-tuned on a narrow vocabulary puts a box on fewer objects outside it, while its in-domain accuracy improves
∙ In-domain accuracy does not determine held-out coverage: across five seeds of one recipe it varies by 1.54 mAP while coverage varies by 2.25 points, and the most accurate seed is the least covered
5 seeds within one recipe; a constructed cross-recipe pair differs by 11.37 points, and a disjoint-support argument covers any in-domain functional
§ 3.2
★ The name goes before the box. After adaptation the model still puts a region on about three quarters of these objects and keeps an eighth of its AP on them, and already at a shallower freeze depth every seed has given up more of its naming than of its boxing
1 arch, the only one scorable by name; 1 domain; all 6 rungs of its ladder, 3 or 5 seeds each but one, every run agreeing at both thresholds and every replicated rung clear of zero
§ 4.1
Which objects it happens to
★ What is damaged is what the checkpoint already had. Categories the pretrained model never learned do not lose coverage; they gain it
3 arch, 752 LVIS categories, 5 seeds; all six cells clear of zero, every seed sharing its cell’s sign
§ 5.3
Table 1: Everything this paper claims, and what each claim was tested on. Ten claims in the four groups the contributions are stated in; ★ marks the five that are new here and ∙ the rest, which confirm or adapt a result the literature already carries (§ 2 ). Every row’s principal measurement rests on five or more independent seeds ; where a row adds a supporting leg carrying fewer, the middle column or the section it names gives the count. The middle column says how widely each claim was tested. Four further claims, resting on less and argued only in the supplement, are in Table .
term
what it denotes
held-out top- K proposal coverage
What we measure. Of the objects whose class the deployment vocabulary never names, the share that at least one of the detector’s top K boxes still lands on (Equation 2 ). Held out means held out of the deployment vocabulary , not of pretraining: the pretrained model was trained on every one of these classes, so this is what a checkpoint keeps , not what it generalises to
class-agnostic localisation
What it stands in for. Putting a box on an object whose name the model was never given. Not measured directly, and coverage can record a miss while this is intact
open-world proposal recall
Where it comes from. Average recall on categories held out of training , an established measurement ( Kim et al, 2022 ; Konan et al, 2022 ) . We keep it and change its use: from scoring a training recipe to reading one checkpoint before and after adaptation
unseen-vocabulary recognition
A different thing. Whether the model can still name categories withheld from adaptation. OWLv2 only, in mAP, never mixed into the coverage axis (§ )
Table 2: Four quantities this paper keeps apart , because they are easy to confuse and only the first is measured here. Table 3 sets it against the instruments that might have been used in its place.
needs
read before the
instrument
a name
class score
categories it scores
read from the detections a model emits
detection AP
yes
no
the ones it names
base/novel split AP ( Wu et al, 2025 )
yes
no
the ones it names
TIDE, Hoiem
yes
no
the ones it names
read from the regions it proposes
Table 3: Why the standard instruments cannot report this loss. Everything above the rule is read from the detections a model finally emits, so it can speak only about categories the model still has a name for. A narrow fine-tune leaves it 22 . Hoiem et al (2012) and Bolya et al (2020) come closest (§ 2 ), but both take their bins after ranking and non-maximum suppression, so a region the model still produces but no longer ranks into its output counts as missed there and as present here (§ 5.1 ). The last two rows read the same quantity of different categories : open-world proposal recall scores what a recipe never learned, Cτ what this checkpoint knew and was not asked about, its top K regions taken by the architecture’s own selection rule.
novel
control
novel,
architecture
condition
trainable
IoU 0.5
IoU 0.7
IoU 0.5
IoU 0.7
unfilt. 0.5
Faster R-CNN
pretrained (COCO)
—
95.10
84.26
94.11
84.00
88.54
tabletop-22, +RPN
14.6M
91.81 ± 0.23
77.32 ± 0.11
95.16 ± 0.11
84.30 ± 0.20
—
tabletop-22, all
41.2M
85.14 ± 1.00
66.58 ± 1.44
85.63 ± 1.11
71.59 ± 1.16
74.96
cardd-22, all
41.2M
78.71 ± 0.51
54.17 ± 0.99
75.83 † ± 1.80
57.82 † ± 1.64
62.78
OWLv2
pretrained
—
97.96
90.06
97.84
91.19
92.99
Table 4: The share of objects each detector still puts a box on, before and after fine-tuning , using its own top 300 regions and no class label. novel is the 21,861 COCO boxes no deployment vocabulary names, control the 2,225 that tabletop-22 does, so control is a naming contrast on the tabletop-22 rows only and † marks every cell where it is not. The last column drops the 1024px2 size floor (Table ). Each ± is one standard deviation over five seeds ; a cell without one is a single run, and --- means not measured, or under trainable that the row trains nothing. Seeds spread widest on OWLv2’s fine-tuned row, where the 95% interval is half the drop it reports. Grounding DINO: Table .
Figure 2: Adaptation damages what the checkpoint already had, and leaves alone what it never had. The 752 LVIS categories split by whether the pretrained model was trained to detect them. Both groups sit outside every deployment vocabulary and are read by the same scorer on the same images, so the split is on pretraining exposure alone. Bars are the mean over five fine-tuned seeds, whiskers a 95% interval; every bar clears zero and every seed shares its bar’s sign (Table ).
pair
domain
raw
partial
both
p
Faster R-CNN / YOLOX
tabletop-22
0.528
0.507
0.594
0.0001
cardd-22
0.626
0.581
0.599
0.0001
Faster R-CNN / RF-DETR
tabletop-22
0.482
0.323
0.324
0.0001
cardd-22
0.536
0.412
0.418
0.0001
YOLOX / RF-DETR
tabletop-22
0.319
0.180
0.191
0.0009
cardd-22
0.389
0.264
0.270
0.0001
Table 5: Different detectors lose the same categories. Each cell correlates two architectures’ per-category drops over the 333 LVIS categories with at least ten boxes, at IoU 0.5 . raw is that correlation; partial first removes each architecture’s own pretrained coverage, and both removes both. p comes from a 10,000 -permutation null that shuffles category labels within one architecture. At IoU 0.7 five of the six cells are larger, the exception being Faster R-CNN against YOLOX on cardd-22 (Table ).
architecture
domain
rungs read
shallowest
worst rung
full
worst is deepest
Faster R-CNN
tabletop-22
5 of 7
−6.94 ( +RPN )
−17.67 ( full )
−17.67
yes
cardd-22
5 of 7
−36.98 ( +RPN )
−36.98 ( +RPN )
−30.08
no
OWLv2
tabletop-22
6 of 6
−0.70 ( heads )
−52.24 ( full )
−52.24
yes
RF-DETR
tabletop-22
6 of 6
−2.01 ( heads )
−10.43 ( dec3 )
−5.12
no
cardd-22
6 of 6
−6.94 ( heads ) †
−43.60 ( dec1 ) †
−12.00
no
YOLOX
tabletop-22
5 of 5
−0.04 ( preds )
−12.87 ( full )
−12.87
yes
Table 6: Freezing more of the network does not reliably protect it : for three of these seven ladders the deepest stage was not the most damaging. Cells are the change in coverage at IoU 0.7 against that architecture’s own pretrained model, so rows are not comparable with one another. rungs read leaves out stages whose trained parameters the objectness score cannot see, which must measure −0.00 (§ ). † marks a rung below 10 in-domain mAP (§ ); those runs are kept here, and dropping them moves two rungs. Five seeds throughout except OWLv2 below full and Faster R-CNN’s interior rungs. Full ladders, and the two rungs: § .
Figure 3: What mixing the pretrained weights back in buys, and what it costs , Faster R-CNN. Each point is one mixing fraction α , running from the fine-tuned model at α=0 to the pretrained one at α=1 ; four fractions carry a marker of their own, keyed above, and the rest are open dots. Up is coverage recovered, left is in-domain accuracy given up, so a curve that rises steeply is buying cheaply. Bars are one standard deviation on each axis. The fractions are not measured on a common set of seeds : five carry α∈{0,0.25,0.5} , three carry {0.2,0.4,0.6} , and the two nearest the pretrained model are single evaluations, marked by a cross. The two domains pay very different prices: tabletop-22 turns almost vertically, cardd-22 trades the two roughly in proportion.