Neural Network Interpretability

Latest papers 203

All topics
CardsList
  1. Is Class Signal Clustered or Routed in Task-Induced Implicit Neural Representation Weight Spaces?

    May 8, 2026Xinyi Guo, Mingyi He, Haobin Ding +7Neural Representation GeometryImplicit Neural Representations

  2. Concept-Based Abductive and Contrastive Explanations for Behaviors of Vision Models

    May 7, 2026Ronaldo Canizales, Divya Gopinath, Corina Păsăreanu +1Abductive ReasoningConcept-Based Explanations

  3. Hyperbolic Concept Bottleneck Models

    May 7, 2026Daniel Uyterlinde, Swasti Shreya Mishra, Pascal MettesConcept Bottleneck ModelsHierarchical Representation Learning

  4. The Metagame of Interpretability and Meta-Attributions

    May 7, 2026Hubert Baniecki, Przemyslaw Biecek, Fabian FumagalliTransformer InterpretabilityFeature Interaction Modeling

  5. AffineLens: Capturing the Continuous Piecewise Affine Functions of Neural Networks

    May 7, 2026Yi Wei, Xuan Qi, Furao Shen +3Neural Network InterpretabilityNeural Network Approximation Theory

  6. Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization

    May 7, 2026Andy Zeyi Liu, Elliot Paquette, John SousRepresentation LearningDecoder-Only Language Models

  7. Interpreting V1 Population Activity via Image-Neural Latent Representation Alignment

    May 5, 2026Xin Wang, Zhuangzhi Gao, Hongyi Qin +3Contrastive LearningNeural Processes

  8. KANs need curvature: penalties for compositional smoothness

    May 4, 2026James BagrowNeural Network Activation FunctionsNeural Network Interpretability

  9. A framework for analyzing concept representations in neural models

    May 2, 2026Burin Naowarat, Hao Tang, Sharon GoldwaterRepresentation DisentanglementNeural Network Interpretability

  10. HyCOP: Hybrid Composition Operators for Interpretable Learning of PDEs

    May 1, 2026Jinpai Zhao, Nishant Panda, Yen Ting Lin +3Neural Surrogate ModelingNeural Network Interpretability

  11. Do Sparse Autoencoders Capture Concept Manifolds?

    Apr 30, 2026Usha Bhalla, Thomas Fel, Can Rager +9Neural Representation GeometrySparse Autoencoders

  12. Towards interpretable AI with quantum annealing feature selection

    Apr 28, 2026Francesco Aldo Venturelli, Emanuele Costa, Sikha O K +3Quantum AnnealingNeural Network Interpretability

  13. SaliencyDecor: Enhancing Neural Network Interpretability through Feature Decorrelation

    Apr 28, 2026Ali Karkehabadi, Jamshid Hassanpour, Houman Homayoun +1Gradient-Based AttributionNeural Network Interpretability

  14. On the explainability of max-plus neural networks

    Apr 27, 2026Ikhlas Enaieh, Olivier Fercoq, García ÁngelFeature AttributionNeural Network Interpretability

  15. Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

    Apr 25, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleyClusteringHierarchical Clustering

  16. Explanation of Dynamic Physical Field Predictions using WassersteinGrad: Application to Autoregressive Weather Forecasting

    Apr 24, 2026Younes Essafouri, Laure Raynaud, Luciano Drozda +1Gradient-Based AttributionNeural Network Interpretability

  17. Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples

    Apr 24, 2026Oussama Bouanani, Jim Berend, Wojciech Samek +2Neural Network Interpretability

  18. H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers

    Apr 23, 2026Ayushi Mehrotra, Dipkamal Bhusal, Michael Clifford +1Gradient-Based AttributionFeature Attribution

  19. Foundation models for discovering robust biomarkers of neurological disorders from dynamic functional connectivity

    Apr 23, 2026Deepank Girish, Yi Hao Chan, Sukrit Gupta +2Medical Imaging Foundation ModelsFunctional Connectivity

  20. An explicit operator explains end-to-end computation in the modern neural networks used for sequence and language modeling

    Apr 22, 2026Anif N. Shikder, Ramit Dey, Sayantan Auddy +6Dynamical SystemsNeural Network Interpretability

  21. Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

    Apr 21, 2026Paweł Pozorski, Jakub Muszyński, Maria GanzhaMultimodal Large Language ModelsNeural Network Interpretability

  22. Decision-Aware Attention Propagation for Vision Transformer Explainability

    Apr 20, 2026Sehyeong Jo, Gangjae Jang, Haesol ParkTransformer InterpretabilityVision Transformer

  23. A Sugeno Integral View of Binarized Neural Network Inference

    Apr 20, 2026Ismaïl Baaj, Henri PradeBinary Neural NetworksNeural Network Interpretability