cs.LGJun 11, 2026

Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance

Authors: Jacques RaynalPierre SlangenElsa RaynalJacques Margerit

Abstract

Learned representations are central to modern machine learning and are commonly evaluated through predictive performance, robustness, uncertainty estimation, and generalization. However, a representation may remain operationally successful while failing to organize persistent residual structures that conventional metrics do not fully capture. This article introduces VER, the Vigilant Evaluator of Representations, a conceptual framework for monitoring representational adequacy. VER does not propose a new learning algorithm, loss function, or model architecture. It defines a diagnostic process for identifying residual structures and assessing whether they may indicate explanatory insufficiency rather than ordinary error, uncertainty, noise, data limitation, or distribution shift. The framework comprises five operations: representation identification, explanatory-domain delimitation, residual-structure detection, explanatory-resistance evaluation, and vigilance signaling. VER distinguishes stable adequacy, a vigilance condition, and a representational alert. It is intended to complement performance evaluation, uncertainty estimation, out-of-distribution detection, robustness analysis, and explainable AI by making representational adequacy an explicit object of inquiry. The article also outlines a path toward empirical evaluation through benchmarks designed to detect representational inadequacy when predictive performance remains satisfactory. VER is conceptual and methodological; it does not prove that a representation is inadequate, select a replacement representation, or provide an operational implementation.

Explore similar work

May 26, 2026cs.AI

RULER: Representation-Level Verification of Machine Unlearning

Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current protocols verify this at the output level through membership inference, retain accuracy, and forget-set accuracy, but a model can satisfy all three whilst still encoding forgotten records in its intermediate representations. We introduce RULER, a set of representation-level verification metrics. The oracle-comparative metric M2 measures whether forget-set records occupy the same representational position as in a model retrained without them. The oracle-free metric M4 detects residuals from the unlearned model's internal similarity structure alone, without retraining. Four approximate unlearning methods all pass output-level evaluation, yet under a linear mixed-effects model M2 detects significant residuals in 10 of 12 conditions (p<0.05), with effect sizes growing as the forget fraction increases. A fifth method, Bad Teacher, shows the same residuals despite a different forgetting mechanism. M4 acts as a pre-unlearning diagnostic across tabular, image, clinical text, and face-identity settings: it detects identity-level memorisation in face recognition models where no tested method fully erases the signal.
Georgina Cosma, Axel Finke
Jun 5, 2026cs.LG

Bootstrap Theory of Representational Emergence: Explanatory Insufficiency as a Driver of Representation Learning and World Models

Representation learning is central to modern machine learning, but most research examines how representations are optimized after a framework has been selected. Less attention is given to when a new representational level becomes necessary. This article introduces the Bootstrap Theory of Representational Emergence (TBER), an initial conceptual theory and research program describing how new representations arise when existing ones become explanatorily insufficient. A representation may remain descriptively useful while failing to make certain observations, relations, transformations, or organizational properties intelligible. TBER treats this explanatory insufficiency as a positive signal for representational transition rather than as simple falsification or prediction error. The proposed recursive process follows five stages: stabilized observation, anomaly detection, recognition of explanatory insufficiency, representational emergence, and provisional stabilization. The framework concerns transitions between scientific or computational representations, not equivalent transitions within the physical systems being observed. Its scope includes representation learning, latent spaces, foundation models, world models, digital twins, adaptive biological systems, and scientific discovery. TBER does not propose a new algorithm, model architecture, benchmark, or optimization procedure. Its contribution is meta-representational: it provides a framework for interpreting when and why new representational levels become necessary. A possible implication for future artificial intelligence is the development of systems able to detect when their own internal representations have reached explanatory limits and to initiate representational refinement or transition.
Jacques Raynal, Pierre Slangen, Elsa Raynal +1
May 19, 2026cs.CV

Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning

Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-level metrics. We challenge these works by introducing Mirage, a representation-level auditing framework that comprises four complementary diagnostics: Linear probe recovery (LPR), centered kernel alignment (CKA), feature separability scoring, and layer-wise recovery analysis. Extensive experiments across seven datasets and seven baseline methods following recent VFL unlearning protocols reveal three key findings: (1) Forgetting gap: methods that pass output-level certification still retain substantial class structure in their representations, with LPR exceeding the retrained baseline by up to 15.4 points; CKA shows that these models remain structurally closer to the original than to the retrained reference, while separability scores indicate persistent geometric discrimination. (2) Unlearning trilemma: no existing method simultaneously achieves high utility, output-level forgetting, and representation-level forgetting. (3) Class-sample asymmetry: class-level forgetting leaves strong representational traces (LPR exceeding 96 percent on several datasets), whereas sample-level forgetting is indistinguishable from chance (LPR is approximately 50 percent); layer-wise analysis further shows that residual class information persists across network depths. These findings call for representation-aware evaluation standards in federated unlearning research. Code is publicly available at https://github.com/YuZhenyuLindy/Mirage.
Zhenyu Yu, Yangchen Zeng, Chunlei Meng +2