Neural Network Interpretability

Latest papers 203

All topics
CardsList
  1. From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models

    Apr 20, 2026Qidong Wang, Junjie Hu, Ming JiangMulti-Task LearningCausal Attribution

  2. Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs

    Apr 20, 2026Charles Ye, Bo Yuan, Lee SharkeyLLM InterpretabilityNeural Network Interpretability

  3. Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models

    Apr 18, 2026Xinru Yan, Boxi Cao, Yaojie Lu +4Unified Multimodal ModelsNeural Network Interpretability

  4. The Query Channel: Information-Theoretic Limits of Masking-Based Explanations

    Apr 17, 2026Erciyes Karakaya, Ozgur ErcetinFeature AttributionExplainable Artificial Intelligence

  5. ProtoTTA: Prototype-Guided Test-Time Adaptation

    Apr 16, 2026Mohammad Mahdi Abootorabi, Parvin Mousavi, Purang Abolmaesumi +1Prototype-Based LearningTest-Time Adaptation

  6. Improving Sparse Autoencoder with Dynamic Attention

    Apr 16, 2026Dongsheng Wang, Jinsen Zhang, Dawei Su +1Sparse AutoencodersNeural Network Interpretability

  7. Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation

    Mar 20, 2026Lasse Marten Jantsch, Dong-Jae Koh, Seonghyeon Lee +1Transformer InterpretabilityLLM Interpretability

  8. Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information

    Mar 18, 2026Shih-Heng Wang, Tiantian Feng, Aditya Kommineni +4Neural Audio CodecsSparse Autoencoders

  9. On the Reliability of Cue Conflict and Beyond

    Mar 11, 2026Pum Jun Kim, Seung-Ah Lee, Seongho Park +2Neural Network Interpretability

  10. Enhancing Physics-Informed Neural Networks with Domain-aware Fourier Features: Towards Improved Performance and Interpretable Results

    Mar 3, 2026Alberto Miño Calero, Luis Salamanca, Konstantinos E. TatsisFourier Feature EmbeddingsNeural Network Interpretability

  11. Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models

    Feb 19, 2026Chenyi Ji, Kian P. Abdolazizi, Hagen Holthusen +2Materials ScienceNeural Network Interpretability

  12. Mechanistic Evidence for Spectral Structures in Prior-Data Fitted Networks

    Jan 29, 2026Kaustubh Sharma, Srijan Tiwari, Ojasva Nema +1Prior-Data Fitted NetworksNeural Network Interpretability

  13. Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI

    Oct 1, 2025Akchunya Chanchal, David A. Kelly, Hana ChocklerExplainable Artificial IntelligenceExplanation Stability

  14. Neural Logic Networks for Interpretable Classification

    Aug 11, 2025Vincent Perreault, Katsumi Inoue, Richard Labib +1Boolean Function LearningNeural Network Interpretability

  15. Position: Use Sparse Autoencoders to Discover Unknowns

    Jun 30, 2025Kenny Peng, Rajiv Movva, Jon Kleinberg +2Explainable Artificial IntelligenceSparse Autoencoders

  16. Learning Interpretable Differentiable Logic Networks for Tabular Regression

    May 29, 2025Chang Yue, Niraj K. JhaEfficient Neural Network InferenceNeural Network Interpretability

  17. Tensorization is a powerful but underexplored tool for compression and interpretability of neural networks

    May 26, 2025Safa Hamreras, Sukhbinder Singh, Román OrúsTensor NetworksNeural Network Interpretability

  18. Explainable embeddings with Distance Explainer

    May 21, 2025Christiaan Meijer, E. G. Patrick BosFeature AttributionPerturbation-Based Feature Attribution

  19. Explainable Bayesian deep learning through input-skip Latent Binary Bayesian Neural Networks

    Mar 13, 2025Eirik Høyheim, Lars Skaaret-Lund, Solve Sæbø +1Bayesian Neural NetworksUncertainty Quantification

  20. Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning

    Feb 2, 2025Zeyu Jiang, Hai Huang, Xingquan ZuoReinforcement LearningConcept-Based Explanations