Neural Network Interpretability

Latest papers 203

All topics
CardsList
  1. The Polytopal Neural Network

    Oct 8, 2026A. Emilie J. Wedenborg, Anders V. Nørskov, Teresa Dorszewski +2Neural Network InterpretabilityGeometric Representation Learning

  2. RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

    Oct 8, 2026Ilya Lasy, Nora Yinuo Cai, Kola AyonrindeMixture of ExpertsExpert Routing

  3. For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance

    Oct 7, 2026Bjørn Leth Møller, Bulat Ibragimov, Christian IgelFeature AttributionNeural Network Interpretability

  4. Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution

    Oct 6, 2026Louis McCallum, Mick GriersonVariational AutoencodersAudio Representation Learning

  5. Interpretable Hypergraph Learning via Neural Additive Models

    Oct 5, 2026Shihan Feng, Xin Zheng, Shiyi Yang +3Neural Network InterpretabilityHypergraph Neural Networks

  6. Weight Oracles: Reading Neural Network Weights with Language Models

    Oct 5, 2026Krishna Kabra, Constantin Venhoff, Christian Schroeder de WittLLM AuditingBackdoor Detection

  7. Disentangling Computation in Multi-Task Neural Networks with the Green's Operator

    Sep 30, 2026James HazeldenRecurrent Neural NetworksNeural Network Interpretability

  8. Synthetic Speech Attribution via Prototypical Networks

    Sep 30, 2026Viola Negroni, Paolo Bestagini, Stefano TubaroNeural Network Interpretability

  9. Certified Approximation for Interpretable Representer Landmarks

    Sep 30, 2026Jayanta Mukherjee, Shourya Verma, Mengbo Wang +2Neural Network InterpretabilityKernel Methods

  10. Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge

    Sep 29, 2026Zixing Jia, Yuhang Pan, Ni JiLink PredictionRelational Reasoning

  11. Explainability from Training with Applications to TCR-Epitope Prediction

    Sep 28, 2026Jiarui Li, Zixiang Yin, Samuel Landry +2Interpretable MLNeural Network Interpretability

  12. Understanding Decision-Making Mechanisms in Neural Routing Solvers

    Sep 28, 2026Fatemeh Askari, Mazdak Teymourian, Mohammad Izadi +1Mechanistic InterpretabilityNeural Network Interpretability

  13. Imprint Reader: From Weight-Update Readout to Behavioral Intervention

    Sep 28, 2026Guanxu Chen, Qihao Lin, Jing ShaoLLM InterpretabilityNeural Network Interpretability

  14. Explaining Hyperbolic Neural Networks via Geometry-Aware Relevance Propagation

    Sep 28, 2026Ping Xiong, Shanglin Li, Yi Ding +2Neural Network InterpretabilityLayer-Wise Relevance Propagation

  15. A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees

    Sep 28, 2026Vojtěch Kůr, Adam Kukučka, Tomáš Brázdil +1Concept-Based ExplanationsNeural Network Interpretability

  16. Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

    Sep 27, 2026Gert Lek, Zixuan Xia, Pin-Yu Chen +1Transformer InterpretabilityFaithfulness of Language Model Explanations

  17. Fisher Simplicity in Kolmogorov-Arnold Networks and Multilayer Perceptrons

    Sep 26, 2026Ami Tavory, Meir FederNeural Network InterpretabilityKolmogorov-Arnold Networks

  18. Learnable Time-Frequency Masks for Explaining Time-Series Classifiers

    Sep 24, 2026Theresa Dahl Frehr, Francisco Pelayo, Lukas Raad +3Feature AttributionTime Series Classification

  19. The Linear Representation Hypothesis Needs a Group Action

    Sep 22, 2026Louie Hong Yao, Yuhao Li, Shengchao LiuLinear Representation HypothesisNeural Network Interpretability

  20. Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability

    Sep 20, 2026Aleks Czufarow, Ihor BabinNeural Network InterpretabilityCross-Architecture Knowledge Distillation

  21. TetrisCNN for interpretable detection of phases of matter from experimental quantum simulator data

    Sep 17, 2026Kacper Cybiński, Björn van Zwol, James Enouen +5Quantum Machine LearningNeural Network Interpretability

  22. NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections

    Sep 15, 2026Alexander Caicedo, Víctor De La Hoz, Santiago AlférezFeature AttributionNeural Network Interpretability

  23. Interpreting hierarchical organisation of speaker embeddings

    Sep 14, 2026Yanze Xu, Wenwu Wang, Mark D. PlumbleySpeech ProcessingNeural Network Interpretability

  24. Prism-SQA: An Interpretable and Adaptable Neural Framework for Surface Electromyography Quality Assessment

    Sep 11, 2026Kuan-Chen Wang, Kai-Chun Liu, Ping-Cheng Yeh +2Neural Network Interpretability

  25. Distributed Lag Neural Additive Models

    Sep 7, 2026Calle Helmersson, Shivang Pandey, Leonardo Olivetti +1Nonparametric RegressionNeural Network Interpretability

  26. Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks

    Sep 1, 2026Osvaldo M Velarde, Lucas C Parra, Alireza Hashemi +1Neural Network InterpretabilityNeural Network Compression

  27. Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures

    Sep 1, 2026Jaee Ponde, Roshni Agarwal, Subhashis BanerjeeNeural Network InterpretabilityConvolutional Neural Networks