cs.CVSep 11, 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

Authors: Kartik Jhawar, Lipo Wang

Abstract

Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain enough information to determine the best action at all? We show that this is not guaranteed, even with a perfect selector. If an observation channel makes two deployments look the same while their TTA rankings differ, reliable selection is impossible from that channel; richer evidence can restore the decision only when it resolves the relevant ambiguity. We make this boundary exact in a finite-batch Gaussian TTA model, where doing nothing beats mean recentering for small shifts, recentering wins beyond a unique critical shift, and the boundary shrinks as 1/n1/\sqrt n. Public benchmark studies on CIFAR-100-C and DomainNet-126 show the same failure mode with modern TTA methods: changing only deployment structure can reverse the oracle action while global order-blind evidence remains unchanged. The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

Explore similar work

CardsList
  1. Dual Strategies for Test-Time Adaptation

    Apr 19, 2026Nam Nguyen Phuong, Duc Nguyen The Minh, Phi Le Nguyen +2Distribution ShiftsDuality