Neural Network Interpretability
Momentum
22 papers in the last four weeks, up 450% on the four weeks before. 0.2% of all new papers.
Latest papers 203
Understanding how deep neural networks process information remains a central challenge. Existing interpretability methods often compromise structural fidelity, rely on prespecified corpora, or explain models post-hoc. We propose Polytopal Neural Networks (PNNs), a framework that extracts distinct layer-wise aspects by enforcing a polytope-based structure that is used directly in subsequent information processing. We scale our approach using learned corpus representations and an amortized simplex inference procedure and highlight how the framework also gives a direct route to vector quantized (VQ) training. In PNNs, observations are explicitly described by their alignment with layer-specific aspects. Empirical results show that imposing polytopal constraints on neural network representations preserves meaningful structures in the latent space with minimal degradation in performance, favorable compressed representations when compared to VQ representations in unsupervised learning, while also providing a performant new approach to VQ deep learning training. Our findings suggest that deep networks can enforce interpretable polytope-based representations, offering a principled path toward more transparent AI systems with minimal performance compromise.
RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.
For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance
The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect the relative importance of input features for a particular prediction. The quality of such maps is often assessed by measuring faithfulness based on the area under insertion and deletion curves, which measures changes in the model output as features are added and removed. In this study, we derive an objective function from this notion of faithfulness and a way to approximate its gradient. We establish the connection between insertion curves and top- feature selection, which leads to a loss function measuring the quality of attributions. Randomization of the loss allows us to efficiently approximate its gradient. To show the effectiveness of the general approach, we combine the loss function with the neural explanation mask framework. The resulting method, termed Ra-NEM, can be used with any differentiable model without affecting the model's performance. Experiments demonstrate that Ra-NEM provides accurate attributions robustly and efficiently. Compared to other algorithms, the attributions have not only higher faithfulness but also perform well in terms of other XAI metrics. The high inference speed of Ra-NEM makes the method suitable for online applications. The code is available online: https://github.com/baerminator/Ra_Nem
Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution
Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
Interpretable Hypergraph Learning via Neural Additive Models
Hypergraphs offer a natural framework for modeling networked data, where dependencies among entities are governed by higher-order interactions. While hypergraph learning methods such as hypergraph neural networks have demonstrated remarkable predictive performance, most existing approaches rely on black-box message-passing architectures, making it difficult to disentangle the contributions of node attributes and higher-order structural information. To address this challenge, we introduce the hypergraph neural additive network (HGNAN), an inherently interpretable framework for learning on hypergraph-structured data. HGNAN extends classical neural additive models to higher-order relational data by integrating feature-wise nonlinear decomposition with hypergraph-aware structural aggregation, enabling transparent prediction for both node- and hyperedge-level tasks. Extensive experiments on benchmark datasets demonstrate that HGNAN achieves performance comparable with state-of-the-art hypergraph learning methods while providing intrinsic and meaningful interpretability.
Weight Oracles: Reading Neural Network Weights with Language Models
Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.
Disentangling Computation in Multi-Task Neural Networks with the Green's Operator
How is computation organized and reused across tasks and time in a trained recurrent network? Most analyses emphasize the geometry of neural activity, dynamical motifs, or local perturbation growth. We instead study the network's global first-order perturbation response. The finite-horizon Green's operator maps perturbations at each source along a trajectory to their downstream state-space responses and therefore directly represents perturbation routing. Simple reductions of this operator provide task-to-task and time-to-time views of the same computation, while matrix-free products make these views accessible without constructing the full operator. In a flexible multitask recurrent network, task reductions reveal structured reuse of known computational motifs, while temporal reductions reveal causal pathways and how they emerge during training. Our main point is simple: the Green's operator provides a global response geometry for mapping the organization of learned dynamical computation.
Synthetic Speech Attribution via Prototypical Networks
Synthetic speech attribution aims to identify the generative system responsible for a speech signal, but current approaches typically rely on black-box neural networks that provide limited insight into their decisions. This work investigates prototype-based networks as an interpretable alternative, where predictions are grounded in comparisons with representative training examples. We adapt ProtoPNet to spectrogram-based speech representations and evaluate the proposed framework on the MLAAD dataset under closed-set, cross-lingual, and open-set conditions. Experiments show that prototype-based reasoning achieves competitive or improved attribution performance compared with the baseline while enabling example-based explanations. These results highlight that interpretability and performance can be jointly achieved in synthetic speech attribution through prototype-based modeling.
Certified Approximation for Interpretable Representer Landmarks
Representer explanations rank the training landmarks that most influence a self-supervised representation. At scale, this ranking rests on up to four stacked approximations of the empirical neural tangent kernel (eNTK). These are random output heads, a parameter sketch, landmark sampling and a coefficient fit. Existing analyses bound each approximation separately, but none certifies the top- set against their combined error. We introduce CAIRN (Certified Approximation for Interpretable Representer laNdmarks), a framework that carries this error through to the ranking. We derive the exact variance of the sketched multi-head eNTK, which matches measurement within where Johnson-Lindenstrauss bounds err by up to . This yields a high-probability top- certificate for a fixed coefficient fit, alongside exact residual-trace certificates for discarded spectral mass. An exact product-variance identity separates kernel error from fit variability and identifies when a larger kernel budget can still sharpen a ranking. Stochastic Lanczos Quadrature (SLQ) estimates the effective dimension within and guides the landmark budget without dense eigendecomposition. We show that residual mass does not control class coverage, and residual-greedy selection cuts the worst coverage excess of -means++ from to ( on the sketched eNTK). Cross-view initializers outperform principal-component initialization in five (AUI) to all six (CSI) settings. Against the KREPES Gauss-Newton solver, CAIRN converges to faster, trails by at most points and gains up to points on MNIST. Together, these results make the reliability of representer explanations measurable and show where approximation budgets are best spent.
Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge
Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representations of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. Across standard knowledge-graph benchmarks, NSR achieves competitive accuracy without leading on every dataset, and has lower reported training times than several neural baselines. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.
Explainability from Training with Applications to TCR-Epitope Prediction
Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR and evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.
Understanding Decision-Making Mechanisms in Neural Routing Solvers
Neural Combinatorial Optimization (NCO) has achieved strong empirical success, yet the internal mechanisms driving model decisions remain largely unexplored. In this paper, we investigate three representative autoregressive NCO models spanning two encoder-decoder configurations: AM and POMO (heavy-encoder, light-decoder), and LEHD (light-encoder, heavy-decoder). Through behavioral analyses, representation probing, and causal interventions, we examine how these models construct solutions and use internal representations during decoding. Our results suggest that AM and POMO predominantly follow a persistent geometric pattern throughout solution construction, whereas LEHD contains linearly accessible information about multiple future actions. Causal experiments further provide evidence for the role of future-node representations in LEHD's decision-making. We also observe that LEHD relies strongly on the current-node representation for immediate local decisions, while the start-node representation plays a broader navigational role over the subsequent route. Cross-instance alignment analyses additionally indicate that LEHD maps current-node representations into a relatively shared latent region, which may provide a stable reference for evaluating subsequent decisions. Across the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem, these results reveal distinct decision-making patterns across these architecturally distinct solvers and provide a foundation for more interpretable analyses of NCO solvers. Code and additional visualizations are provided in the https://github.com/NCO-Interpretability/NCO-Interpretability.
Imprint Reader: From Weight-Update Readout to Behavioral Intervention
As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter-level traces that can, in principle, be inspected directly. However, current models cannot decode these traces into an explicit account of what they have learned. To this end, we introduce the \textit{Imprint Reader}, a model trained with \textit{Semantic Mount-and-Read Tuning} (SaRT) to describe frozen weight updates. SMaRT mounts each update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, while no-change and random-perturbation controls discourage unsupported claims. On held-out updates, the joint Reader reaches judge-based Pass@100 of for knowledge and for behavior. These results demonstrate the feasibility of natural-language readout while pointing to reliability across updates as the next step. Beyond free-form generation, the Reader provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update. Its coordinate-aligned gradients support intervention through MetaEdit. At a pruning rate, Reader-guided selection raises measured harmful-prompt refusal from to under a safety-maintenance target. Using behavior descriptions without target-task training data, MetaEdit increases the frequency of backtracking and sub-goal expressions in mathematical reasoning traces and raises BFCL Overall from to .
Explaining Hyperbolic Neural Networks via Geometry-Aware Relevance Propagation
Hyperbolic neural networks introduce geometric operations that require explicit treatment in relevance propagation. Equivalent geometric realizations can produce different feature attributions, even when local relevance is conserved. We study this problem through Geometric Representation Invariance (GRI), a specialization of Implementation Invariance, and zero-curvature consistency, which requires identity relevance propagation when a geometric module approaches the identity. We propose LRP-radial-all for origin-centered radial modules, treating geometric scaling as modulation and assigning relevance entirely to the signal branch. The rule conserves relevance, is invariant to equivalent radial factorizations, and satisfies zero-curvature consistency, yielding GRI for a specified Poincaré-Lorentz logarithmic-map construction. In contrast, a conservative LRP-half baseline can violate both consistency criteria. Experiments on hyperbolic MNIST, sEEG, and CIFAR-10 classifiers assess attribution fidelity, qualitative explanations, and runtime. LRP-radial-all achieves competitive attribution fidelity across datasets with runtime comparable to GradientInput and substantially lower than Integrated Gradients. These findings motivate geometry-aware propagation rules that distinguish relevance conservation from consistency across equivalent computations.
A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning
Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the first time into a specifically designed variational autoencoder (VAE) structure, to solve a general class of nonlinear PDEs on smooth manifolds. AFDONet is the first neural PDE solver whose architectural and component design is fully guided by an established mathematical framework (in this case, AFD theory), turning neural operator design from an art to a science. Thus, AFDONet also exhibits exceptional mathematical explainability and groundness, and enjoys several desired properties. Furthermore, AFDONet achieves outstanding solution accuracy and competitive computational efficiency in several benchmark problems. In particular, thanks to its deep connections with AFD theory, AFDONet shows superior performance in solving PDEs on i) arbitrary (Riemannian) manifolds, and ii) datasets with sharp gradients. Overall, this work presents a new paradigm for designing explainable neural operator frameworks.
Understanding Confabulation and Rethinking Reconstruction in Activation Explanations
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Fisher Simplicity in Kolmogorov-Arnold Networks and Multilayer Perceptrons
Kolmogorov-Arnold Networks (KANs) are motivated in part by interpretability: their learned edge functions can be inspected, pruned, and reduced to symbolic structure. In a fixed-basis KAN, this makes a small or zero basis coefficient look like a certificate of simplicity, much as a dead rectified linear unit (ReLU) marks unused computation in a multilayer perceptron (MLP). Fisher nullity gives a precise statistical notion: a parameter direction is Fisher-simple exactly when perturbing it is invisible under the task distribution. We study when these architectural and Fisher notions agree. For a dead ReLU unit, they agree: the closed activation region makes the associated score directions vanish. For a fixed-basis KAN, they do not. In the single-layer Gaussian case, the coefficient Fisher matrix is a basis Gram matrix under the input distribution and is independent of the fitted coefficients. In a multilayer KAN, Fisher simplicity is graph-path based: the data must reach a basis atom and its perturbation must propagate through the downstream network. We encode these two conditions in an effective edge measure and, under local dictionary independence and effective-measure nondegeneracy, show that zero effective exposure exactly identifies Fisher-null directions within an edge. Controlled diagnostics confirm that zero coefficients can preserve rank while effective path disconnections remove the predicted directions. Coefficient magnitude alone is therefore not a Fisher-based pruning criterion for KANs.
Learnable Time-Frequency Masks for Explaining Time-Series Classifiers
Time-series explainability remains challenging because discriminative information is often encoded in latent frequency or time-frequency features rather than in the raw signal itself. Existing attribution methods typically operate either in the time domain or in a fixed transform domain, limiting their ability to capture salient information across different representations. We propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms. We evaluate the framework on the STFT, the continuous wavelet transform, and the discrete wavelet transform. In addition, we extend the virtual inspection layer approach from the STFT to both wavelet transforms, enabling LRP to generate explanations in these representations. On a synthetic dataset, XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines. Across two real-world datasets, XACT produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria. These results demonstrate that learning explanations directly in time-frequency representations offers a flexible approach to interpreting deep-learning models for time series data.
The Linear Representation Hypothesis Needs a Group Action
To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.
Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability
Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical in safety-critical domains such as autonomous driving and medicine. This study investigates whether Knowledge Distillation transfers the spatial feature attribution of a large teacher network to a compact student. To assess the influence of the KD scheme on interpretability, we distill a ResNet-152 teacher into a ResNet-34 student on ImageNet-1K across five configurations by systematically varying the distillation temperature and soft-label loss weight. Models are evaluated on top-1 accuracy, along with two interpretability metrics: Relevance Mass Accuracy and Relevance Rank Accuracy. These metrics are computed via Grad-CAM heatmaps benchmarked against ground-truth object masks. Our results show that top-1 accuracy ranges from 71.6% to 74.0%. For Grad-CAM, RMA ranges from 7.7% to 9.7% and RRA from 7.3% to 10.1%; for Guided Grad-CAM, RMA ranges from 16.1% to 18.6% and RRA from 15.9% to 21.5%. Interpretability proves far more sensitive to the soft-label weight than to the temperature: keeping the student anchored to hard labels preserves both accuracy and coarse localization, whereas weighting the teacher heavily degrades both. Fine-grained attribution, however, fell below the undistilled baseline in every configuration tested, indicating that logit distillation transmits where a model attends more readily than the pixel-level structure of that attention. We evaluate 12 cross-architecture combinations of convolutional and transformer-based models, revealing that the inheritance of fine-grained spatial reasoning is fundamentally bottlenecked by the student's intrinsic structural biases. To our knowledge, this is the first application of this interpretability-aware evaluation framework - previously used for neural network pruning - to KD.
TetrisCNN for interpretable detection of phases of matter from experimental quantum simulator data
Detecting phases of matter in general relies on identifying the correct order parameter - a task that remains notoriously difficult for unknown transitions and traditionally is guided by physical intuition and educated guess. Neural networks have recently offered an alternative route by locating phase transitions in known models without any a priori physical knowledge. Yet these approaches remain black boxes and only identify phases without elucidating their properties. Moreover, they often struggle when confronted with realistic, noisy experimental data, which constitute the ultimate testbed for automated methods in physics. Here, we bridge these perspectives by introducing TetrisCNN, a convolutional architecture with parallel branches of differently shaped filters, reminiscent of Tetris blocks, that learns sparse, interpretable latent representations directly in terms of spin correlators. Applied to experimental snapshots of two-dimensional Ising and XY quantum simulators measured in multiple bases, the network not only detects phase transitions and crossovers but also expresses its latent representation and decision boundaries as symbolic formulas built from experimentally measurable spin correlators. This framework opens the way to integrating interpretable neural networks with quantum simulators to uncover and understand new phases of matter.
NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections
Understanding how deep neural networks make decisions remains a fundamental challenge. We present NObSP (Nonlinear Oblique Subspace Projections), a framework that decomposes predictions into explicit per feature contribution functions and an interaction residual. NObSP exploits the linear final layer of a trained network and uses oblique projections in sample space to reduce double counting when learned feature subspaces overlap, thereby supporting both local explanations and global functional analysis. We establish connections to functional ANOVA and the Kolmogorov-Arnold representation theorem and derive an efficient partial regression algorithm for out of sample evaluation. For convolutional networks, NObSP-CAM produces class activation maps without backward passes after a one time calibration. Experiments on tabular and vision benchmarks show faithfulness comparable to established attribution methods. On a synthetic benchmark with known component functions, NObSP obtains a Function Reproduction Score of 0.989, compared with 0.966 for KernelSHAP and 0.922 for Integrated Gradients. On TinyImageNet, contribution vector embeddings improve mean nearest neighbor class purity from 0.654 for raw activations to 0.713 and reduce mean neighbor distance by more than half. These results indicate that NObSP complements scalar attribution methods by recovering functional contribution profiles with separable positive and negative evidence.
Interpreting hierarchical organisation of speaker embeddings
Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utterances to recognise speaker identities. However, the internal mechanisms of these networks remain largely opaque, motivating research in explainable artificial intelligence (XAI) to understand them. Nevertheless, existing studies have analysed how speaker embeddings are organised, but rarely frame these analyses within XAI. Hence, this work proposes to explain and interpret the organisation of speaker embeddings from an XAI perspective. To this end, we apply a hierarchical clustering algorithm, Single-Linkage Clustering (SLINK), to analyse whether some speaker embeddings naturally form clusters with hierarchical relationships. The resulting hierarchical organisation (i.e. hierarchical clusters) is evaluated using the Cluster-Class Matching (CCM) method. Moreover, we propose a new method, termed Hierarchical Cluster-Class Matching (HCCM), to identify which hierarchical clusters best match individual semantic classes (e.g. male) and conjunctive semantic classes (e.g. UK & male), thereby interpreting the clusters using their matched classes. The matching degree is quantified using a new metric called the L-score, which makes imperfect matches diagnosable. HCCM's results show that hierarchical clusters analysed by SLINK are interpreted using different classes related to speaker identity, gender, and nationality, providing insight into semantics within the hierarchical organisation of our examined speaker embeddings.
Interpreting the predictions of neural network classification based on a Taylor Coefficient Analysis (TCA)
We introduce a rigid and comprehensive taxonomy and paradigm for characterizing the influence of the input feature space on the predictions of a neural network (NN) used for event classification, based on a Taylor expansion of in . The complete process of introspection we refer to as Taylor Coefficient Analysis (TCA). Based on two simplistic example tasks, which can be easily understood and bencmarked, we illustrate the power of the TCA when it comes to revealing, what properties of have led to what value of , of a given NN model, building up intuition for the method. A more complex application is meant to represent of a typical classification task at a CERN LHC experiment. Based on this application, we play through the different levels of introspection that the TCA offers and discuss a number of practical aspects for a TCA application typical for the analysis of CERN LHC data. We conclude with a study to support the assumption that those properties of most relevant for tasks of the complexity typical for a CERN LHC experiment, are usually caught by a TCA up to the second order.
Prism-SQA: An Interpretable and Adaptable Neural Framework for Surface Electromyography Quality Assessment
sEMG is vulnerable to various contaminants that distort signal morphology and spectral content. Accurate signal quality assessment (SQA) is essential for identifying such degradation and ensuring reliable clinical analyses and decisions. Recent neural network-based SQA methods achieve accurate quality estimation by learning complex contamination patterns, yet their black-box nature prevents clinicians from understanding or validating the reported scores and limits adaptability to application-specific quality definitions without retraining. To address these limitations, we propose Prism-SQA, an interpretable and adaptable neural framework that reformulates SQA as a physiology-aware source-separation and verification process. Prism-SQA decomposes each input signal into a clean sEMG component and five contaminant-specific components using a U-Net with bidirectional long short-term memory. Each separated contaminant component is examined by a Contaminant Fingerprint Verifier, which enforces physiological plausibility by comparing its temporal and spectral structure with canonical contaminant signatures. This design allows clinicians to inspect how each contaminant affects signal quality, grounding the assessment in transparent, signal-level evidence rather than opaque latent representations. Quality indices computed from the verified components further enable customization of quality criteria across clinical contexts without retraining. We evaluate Prism-SQA on continuous quality-score estimation using synthesized noisy sEMG from public Ninapro datasets and on binary quality classification using a clinical dysphagia dataset. Results show that Prism-SQA achieves competitive or better performance than contemporary black-box neural methods while providing explicit interpretability and adaptability, advancing toward practical and clinically aligned sEMG SQA.
Interpreting Object-Dependent Concept Brittleness in Text-to-Image Diffusion Models
Although text-to-image diffusion models generally exhibit strong prompt-following ability, we identify a persistent and previously underexplored failure pattern in which a small subset of prompts differing only in the object consistently fails to realize the same target concept under identical generation settings. We term this phenomenon object-dependent concept brittleness. Such cases suggest systematic internal blind spots rather than random sampling noise. In this paper, we present an interpretability-oriented framework to audit and minimally correct these failures. Our key idea is to analyze denoising trajectories in a step-wise sparse autoencoder (SAE) space, where abstract style and attribute concepts become more separable than in the raw denoising representation. This sparse space enables us to compare successful and failed generations, identify concept dimensions whose evidence is missing, weakened, or temporally delayed, and construct class-level concept prototypes from reliable class-consistent samples. Based on this audit process, we introduce a lightweight inference-time correction strategy that interpolates denoising features toward the corresponding prototype in SAE space. Rather than serving as a task-specific retraining method, this intervention acts as a validation of the diagnosed concept deficiency. We evaluate the proposed framework on style and attribute failure cases across multiple diffusion backbones, with significant improvements in concept consistency, text fidelity, and repair success. Further analyses show that deeper denoising representations provide clearer concept structure, while early-stage intervention offers the strongest correction leverage. Code is available at https://github.com/Metecade/Object-Dependent-Concept-Brittleness.
Distributed Lag Neural Additive Models
We introduce Distributed Lag Neural Additive Models (DLNAMs), neural-additive analogues of Distributed Lag Non-linear Models (DLNMs) for learning nonlinear effects distributed over lags. DLNAMs replace a prespecified spline cross-basis with neural components that learn exposure--lag response surfaces, avoiding choices of basis family, dimension, and knot placement while preserving additive interpretability and familiar distributed-lag summaries. Exp-centered input layers, smooth activations, and learned subnetwork mixtures produce smooth, locally adaptive representations; pointwise uncertainty combines a conditional last-layer Laplace approximation with between-member ensemble variation. In simulations, DLNAMs generally outperformed DLNM comparators, including penalized and treed variants, in recovering known response functions, with lower bias, stronger boundary recovery, and better-calibrated cumulative intervals; gains were largest for more demanding functions. The architecture performed consistently across sample sizes, outcome families, lag horizons, and jointly fitted multi-exposure settings, retaining recovery performance as exposures were added; fit-specific changes were largely confined to optimization, and applications recovered established empirical patterns.
Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks
Artificial neural networks are often regarded as powerful yet opaque black boxes. Here, we demonstrate that learning in deep neural networks generates local symmetries known in graph theory as fibrations and coverings. We prove that covering symmetries are stable attractors of stochastic gradient descent. Consistent with this theory, we report the emergence of covering symmetries across major network architectures, including multilayer, convolutional, recurrent, and transformer networks. Exploiting these symmetries enables drastic model compression - reducing networks to 17% of their original size without sacrificing performance. Furthermore, controlled breaking of covering symmetry overcomes the loss of plasticity, achieving state-of-the-art performance in continual learning. The theoretical results provide a new foundation for AI systems based on symmetries that convert black boxes into interpretable colored graphs and enable more efficient inference and lifelong learning.
Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures
Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented internally. Our study asks if these networks exhibit \textit{conceptual separation}: if examples of the same concept form coherent representations, and whether related concepts lie closer together in the representation space. We examine this conceptual organisation in Convolutional Neural Networks (CNNs) and Large Language Models (LLMs) through geometric and distributional analysis of their internal activations. In CNNs, familiar ImageNet concepts form coherent and semantically ordered representations, while this coherence weakens for unseen concepts and suffers within-class domain shift. In LLMs, clearly distinct domains remain well separated, related subdomains move closer together, and the distinction between ambiguous topics collapses at both the mean and covariance level. These results suggest that conceptual separation can reveal structure that output accuracy alone cannot, and may serve as a useful diagnostic of how robustly a model represents the concepts it is asked to identify. Code and data available on GitHub.