From internal representations to model improvement through prediction errors
Organizations: Adansons Corp., 214 Homer Ave, Palo Alto, 94301, CA, USA. · Tohoku University, 6-6-11-803 Aza-Aoba Aramaki, Aoba-ku, Sendai, 980-8579, Miyagi, Japan.
Abstract
With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but those signals do not directly capture changes in the model being improved. The target model's own internal features reflect what it has learned so far and change with retraining, making them a natural cue for choosing the next training data. However, feature rarity alone does not reveal the errors that matter for performance. Here we link internal features to prediction errors and their expected impact on performance and select images for labeling and retraining without using labels for candidate images. We evaluated the method with an object detector on two datasets and two pairs of random seeds. Adding internal features improved the identification of prediction errors in 15 of 16 conditions. When performance was averaged over successive labeling rounds, the method outperformed selection based only on feature rarity in all four evaluation settings and ranked among the top two of six methods. With other conditions held fixed, performance after retraining was again higher than with rarity-based selection, even though the latter collected more errors. With longer retraining, the proposed method ranked first among six methods. These results suggest that linking a model's internal features to its errors and their effects on performance may help select training images that improve performance, thereby allowing the model's current state to guide which images are labeled next.
Figures & tables
| Dataset | Seeds | Pro- posed | Ran- dom | External -center | BADGE | Output entropy | Rarity only | Rank |
|---|---|---|---|---|---|---|---|---|
| BDD100K | 11/1 | 13.929 | 13.308 | 13.829 | 12.803 | 12.809 | 12.881 | 1st |
| BDD100K | 18/0 | 13.248 | 13.182 | 13.707 | 12.736 | 12.298 | 12.134 | 2nd |
| COCO 2017 | 11/1 | 15.839 | 15.959 | 14.707 | 15.664 | 15.484 | 15.228 | 2nd |
| COCO 2017 | 18/0 | 15.735 | 15.541 | 14.346 | 15.336 | 14.887 | 14.850 | 1st |
| Method | Information used | Selection rule and role |
|---|---|---|
| Random sampling | Candidate image IDs and the image-selection seed | Draws uniformly at random from the candidate set. Serves as the reference set before the proposed method replaces images. |
| -center selection on features from a separate model | Image features from a ResNet-50 that is separate from the detector under training | Normalizes the 2,048-dimensional features of a ResNet-50 pretrained on ImageNet-1K and repeatedly selects the image farthest in feature space from the labeled set. Uses neither the outputs nor the internal features of the target detector. |
| Output entropy | Final outputs of the target detector | Aggregates, per image, the binary entropy computed from the top-class confidence of each predicted box. Uses neither internal features nor calibrated error probabilities. |
| Selection by rarity of the internal representation | Internal features of the target detector | Computes the minimum cosine distance between each image’s internal representation and those of the images labeled so far, and selects the most distant images. Uses neither output probabilities, error types, nor error probabilities. |
| BADGE adapted to object detection | Internal features and top-class confidence of the target detector | L2-normalizes the mean internal feature of each image, scales it by 1 minus the maximum box confidence, and draws the batch by -means++ seeding on these vectors. Retains the uncertainty and diversity components of BADGE without loss gradients (Methods). Compared as the pre-specified representative of existing methods that use the internal representation. |
| Proposed method | Internal features, outputs, box geometry and neighbor information of the target detector, and a separate 1,500-image labeled set for error estimation | Estimates the error type and probability of each box in the unlabeled candidates and computes an image priority from the effect of each error type, the sufficiency of the evidence and the estimated annotation effort. Starts from the same random draw and replaces, under constraints, only images whose priority difference is confirmed by the stability bounds. |