Neural Network Interpretability

Latest papers 203

All topics
CardsList
  1. A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution

    Jul 10, 2026Jose Luis Lima de Jesus SilvaNeural Algorithmic ReasoningNeural Network Interpretability

  2. All you need is SAMPAT

    Jul 10, 2026Jayadeva, Madhur AswaniShallow Neural NetworksInterpretable ML

  3. How are linear representations learned? Exact solutions to the dynamics of abstraction

    Jul 9, 2026William W. Yang, Andrew M. Saxe, Peter E. LathamLinear Representation HypothesisDeep Linear Networks

  4. Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence

    Jul 9, 2026Yann Claes, Pierre Geurts, Vân Anh Huynh-ThuInterpretable MLNeural Network Interpretability

  5. Naming the Concepts Classifiers Rely On: Language-Anchored Decomposition for Faithful Explanation

    Jul 8, 2026Ahsan Habib Akash, Dipkamal Bhusal, Stacey Jones +3Concept-Based ExplanationsNeural Network Interpretability

  6. Unraveling Machine Behavior by Multi-Level Bias Analysis and Detection: Methodology and Application to Computer Vision

    Jul 8, 2026Ignacio Serna, Aythami Morales, Julian FierrezNeural Network InterpretabilityAlgorithmic Bias

  7. Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories

    Jul 8, 2026Tobias Göbel, Julian R. Ebelt, Zier Mensch +2Phase TransitionsNeural Network Interpretability

  8. On the Principles of Deep Feedforward ReLU Networks

    Jul 8, 2026Changcun HuangReLU Neural NetworksNeural Network Interpretability

  9. X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models

    Jul 7, 2026Jie Huang, Pengfei Yin, Zihan Xu +3Neural Surrogate ModelingNeural Network Interpretability

  10. Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders

    Jul 5, 2026Sohaib AfifiGradient-Based AttributionCounterfactual Explanations

  11. Individual Parameters in Weight-Sparse Transformers Appear Interpretable

    Jul 3, 2026Arnau Marin-Llobet, Stefan HeimersheimTransformer InterpretabilityTransformer

  12. Self-explainable Operator Learning for Discovering Spatial Patterns in Functional Data

    Jul 2, 2026Mojgan Alishiri, Amirhossein ArzaniNeural Network InterpretabilityOperator Learning

  13. Geometry-Aware R-Structured Kolmogorov-Arnold Networks

    Jul 1, 2026Sergei Kucherenko, Nilay ShahNeural Network InterpretabilityKolmogorov-Arnold Networks

  14. Neuro-Bayesian-Symbolic Residual Attention Shallow Network: Explainable Deep Learning for Cybersecurity Risk Assessment

    Jun 29, 2026Nicolaie Popescu-Bodorin, Madeleine TogherShallow Neural NetworksExplainable Artificial Intelligence

  15. Improved Predictive Performance and Interpretability for Mesomorphic Neural Networks Using Local Fidelity Regularization

    Jun 29, 2026Hugo L. Hammer, Vajira Thambawita, Kristoffer Herland Hellton +1Interpretable MLNeural Network Interpretability

  16. Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

    Jun 25, 2026Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. HosseiniActivation SparsitySparse Autoencoders

  17. Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis

    Jun 24, 2026Felipe Moreno, Sharifa Alghowinem, Hae Won Park +1Depression DetectionNeural Network Interpretability

  18. Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds

    Jun 23, 2026Thomas Fel, Matthew Kowal, Mozes Jacobs +22Neural Representation GeometryActivation Sparsity

  19. Interpretable Kolmogorov-Arnold Network with Feature-Isolated Temporal Attention Mechanism for Electricity Load Forecasting

    Jun 22, 2026Jinhao Li, Hao WangMultivariate Time Series ForecastingTime Series Forecasting

  20. Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

    Jun 22, 2026Yanze Xu, Mark D. Plumbley, Wenwu WangGradient-Based AttributionFeature Attribution

  21. From Handcrafted Features to Functional Edge Learning: Evolution of EEG Seizure Detection Frameworks

    Jun 20, 2026Sepideh Kheirollahi, Mohammad Rasoul RoshanshahElectroencephalographyEEG Seizure Detection

  22. What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study

    Jun 20, 2026Yaozhong Kang, Jiang Wang, Runwu Shi +3Direction-of-Arrival EstimationNeural Network Interpretability

  23. What Do Lorentz-Equivariant Jet Taggers Learn?

    Jun 19, 2026Jay Agarwal, Siddharth Khare, Dhruv KumarJet TaggingEquivariant Neural Networks

  24. LISE : Listenable Interpretable Speaker Embeddings

    Jun 19, 2026Xiaoliang Wu, Chongxin Gan, Ke Liu +2Audio Representation LearningSpeaker Verification

  25. A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning

    Jun 19, 2026Talha Ilyas, Deval Mehta, Zongyuan GeConcept-Based ExplanationsNeural Network Interpretability

  26. Neural Additive and Basis Models with Feature Selection and Interactions

    Jun 18, 2026Yasutoshi Kishimoto, Kota Yamanishi, Takuya Matsuda +1Feature Interaction ModelingNeural Network Interpretability

  27. Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks

    Jun 18, 2026Haocheng Duan, Yuxin Guo, Jieyi Bi +4Neural Network InterpretabilityNeural Combinatorial Optimization