Neural Network Interpretability

Latest papers 203

All topics
CardsList
  1. ICON Decomposition: Auditing deep neural networks for shortcuts by decomposing layer-wise representations using concepts

    Aug 26, 2026Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer +7Concept-Based ExplanationsNeural Network Interpretability

  2. Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

    Aug 13, 2026Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8Transformer InterpretabilityLLM Interpretability

  3. The Neural Division of Labor: Biologically-Inspired Modular Architectures for Robust Neuromorphic Computing

    Aug 8, 2026Maksim Bazhenov, Serafim Grubas, Vakhtang PutkaradzeEfficient Neural Network InferenceNeural Network Interpretability

  4. Tools to Explain Neural Networks for Power System Dynamics

    Aug 8, 2026Petros Ellinas, Johanna Vorwerk, Spyros ChatzivasileiadisNeural Surrogate ModelingNeural Network Interpretability

  5. The Spectral Neuron

    Aug 8, 2026Alex ShtoffNeural Network RobustnessNeural Network Interpretability

  6. Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks

    Aug 7, 2026Quanshi Zhang, Qihan Ren, Siyu LouNeural Network GeneralizationNeural Network Interpretability

  7. The Neural Echo: A Signal Processing Perspective for Understanding Neural Networks

    Aug 5, 2026Chongbiao Wang, Daniel Gaa, Joachim Weickert +1Neural Network InterpretabilityConvolutional Neural Networks

  8. Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

    Jul 31, 2026Yifan Wang, Jinyi Mu, Mayank Jobanputra +4Reward ModelingNeural Network Interpretability

  9. Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification

    Jul 31, 2026Stanislaw Janik, Michal ByraImplicit Neural RepresentationsNeural Network Interpretability

  10. Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?

    Jul 28, 2026Qitao Chen, Dongfu Yin, F. Richard YuNeural Network InterpretabilityWeight-Space Analysis

  11. LAWFUL: Law-Aligned Witness for Faithful Use of Latents

    Jul 26, 2026Kevin Chen, Kenneth W. Parker, Anish AroraTransformer InterpretabilityMechanistic Interpretability

  12. Variable Importance Identification Through Lazy Training for Binary Classification

    Jul 25, 2026Anand Singh, Luke Pennella, Eshan Kabir +1Feature AttributionNeural Network Interpretability

  13. Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations

    Jul 24, 2026Kazi Kamruzzaman Rabbi, Md. Zami Al Zunaed Farabe, M. Sohel RahmanFeature AttributionNeural Network Interpretability

  14. Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification

    Jul 23, 2026Guillermo Rubiños Rodríguez, Martín Ottavianelli, Mateo Alonso +4Quantum Convolutional Neural NetworksQuantum-Inspired ML

  15. Neural Feature Governance: Extending Atom Prevalence

    Jul 23, 2026Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait FokouéBayesian Neural NetworksUncertainty Quantification

  16. Scaling Interpretable Transformers with Parity Bottleneck Layers

    Jul 22, 2026Andrew Mack, Kraig Yuheng Tou, Mark Henry +2Feature SuperpositionTransformer Interpretability

  17. Do Sheaf Neural Networks Use Holonomy? A Measure--Intervene--Control Study

    Jul 21, 2026Ankit Grover, Rémi BourgerieNeural Network InterpretabilitySheaf Neural Networks

  18. Tensor Network Machine Learning for Wildfire Susceptibility Mapping: from Grokking Dynamics to Quantum Mixedness of Class Representations

    Jul 21, 2026Domenico Pomarico, Alessandra Costantino, Gabriel Ramirez Sanchez +12Tensor NetworksNeural Network Interpretability

  19. Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent

    Jul 17, 2026Sriram Balasubramanian, Soheil FeiziMechanistic InterpretabilityNeural Network Interpretability

  20. Kolmogorov--Arnold Networks for Small Language Models

    Jul 17, 2026Felippe Alves, Renato VicenteSmall Language ModelsNeural Network Interpretability

  21. TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation

    Jul 16, 2026Wen Yang Tan, Jiawei Li, Fang Liu +4RUL EstimationNeural Network Interpretability

  22. Understanding Structured Health Data through Interaction-Aware Mixture-of-Experts

    Jul 14, 2026Ji Hwan Park, Ying Ding, Tianjin GuoMulti-View LearningClinical Outcome Prediction

  23. From Neural Network Decisions to Training Cases: An Exact Account via Case-Based Decision Theory

    Jul 13, 2026Manli Yan, Yuebin Lin, Yaowen Yu +1Algorithmic AuditingNeural Network Interpretability