Interpretability

Latest papers 309

All topics
CardsList
  1. Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability

    Oct 1, 2026Roberto Stanzione, Jules Barbe, Magali Parrino +2Time-Series Anomaly DetectionTime Series

  2. Interpreting Reasoning of Large Language Models via Partial Information Decomposition

    Sep 30, 2026Barproda Halder, Qiuyi Zhang, Sanghamitra DuttaLarge Reasoning ModelsReasoning Trajectory

  3. Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations

    Sep 29, 2026Hanwei Zhang, Tianma Hu, Gaojie Jin +2Dermoscopic Concept Bottleneck ModelsRobustness Verification

  4. Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing

    Sep 29, 2026Rishabh Mondal, Nipun Batra, Utkarsh MallRemote SensingEarth Observation

  5. Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

    Sep 29, 2026Zhenting Huang, Bo Jiang, Junnan Liu +2Improving Sparse AutoencodersTop-K

  6. CHOQOLATE: Organizing Concept Bottleneck Latent Spaces with Choquet Integrals

    Sep 29, 2026Rémi Kazmierczak, Johanne Cohen, Marianne ClauselConcept Bottleneck ModelsInterpretability

  7. Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving

    Sep 29, 2026Katharina Winter, Stefan Englmeier, Fabian B. FlohrAutonomous DrivingBlind Spots

  8. Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge

    Sep 29, 2026Zixing Jia, Yuhang Pan, Ni JiBrain-Inspired Synergistic FrameworkNeural Representations

  9. Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions

    Sep 29, 2026Hong Je-Gal, Hyun-Suk LeeSchedulersSchedule

  10. Verifying the Linear Representation Hypothesis: How Interpretable Are Vision SAEs?

    Sep 28, 2026Teodor Chiaburu, Franz Motzkus, Frank Haußer +1Improving Sparse AutoencodersInterpretability

  11. SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making

    Sep 27, 2026Shyam Sundar Murali Krishnan, Dean Frederick HougenInterpretabilityBlack Box

  12. Augmenting Visual Anomaly Detection with Automated Interpretability

    Sep 27, 2026Antonio De Santis, Arsenio Leo, Marco BrambillaMultimodal Anomaly DetectionUnsupervised Detection

  13. Grammatical "grandmother neurons" are rare in LLMs

    Sep 24, 2026Linyang He, Nima MesgaraniLinguisticsNeurons

  14. Script Choice in LLMs: Evidence for Late-Layer Commitment

    Sep 23, 2026David Kletz, Sandra Mitrović, Itay Sabato +2Cryptographic Merkle Tree CommitmentsCommitment

  15. The Linear Representation Hypothesis Needs a Group Action

    Sep 22, 2026Louie Hong Yao, Yuhao Li, Shengchao LiuRepresentation SpaceInterpretability

  16. High-Dimensional Online Change Point Detection with Adaptive Thresholding and Interpretability

    Sep 21, 2026Sven Jacob, Bardh Prenkaj, Weijia Shao +1Change-Point DetectionAdaptive Thresholding

  17. G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation

    Sep 20, 2026Yixuan Liang, William Chen, Yunan Wang +4Category-Level Object Pose EstimationRobotic Manipulation

  18. Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability

    Sep 20, 2026Aleks Czufarow, Ihor BabinKnowledge DistillationInterpretability

  19. Flag Game: A Toy Model for Mechanistic Swarm Interpretability

    Sep 16, 2026Elizabeth Pavlova, Hidenori TanakaCollective BehaviorsSwarms

  20. Using OCR Heads to Verbalize Image Semantics

    Sep 16, 2026Sheridan Feucht, Benno Krojer, Sarah Wang +3Optical Character RecognitionSemantic Representations

  21. A unified framework for global and local interpretability using adaptive derivative-ordered random explanation

    Sep 15, 2026Lemen Chao, Ming Lei, Anran FangInterpretabilityFeature Importance

  22. Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

    Sep 15, 2026Younes Boufouss, Luc Pommeret, Thomas Gerald +2Natural Language InferenceInterpretability

  23. Learned Look-Ahead Splitting Rule for CART

    Sep 14, 2026Andrew Gao, Tianlin Liu, Ruichen Han +1Decision TreesLookahead

  24. Data storytelling meets interpretable machine learning: Decoding AI decisions for non-experts without revealing sensitive data and model details

    Sep 14, 2026Lemen Chao, Zixuan Yang, Anran Fang +2InterpretabilityNarratives

  25. The Misery of Mechanistic Interpretability: A Formal Perspective

    Sep 14, 2026Tobias Ladner, Matthias AlthoffMechanistic InterpretabilityInterpretability

  26. Data Attribution at Scale via Influence Matrix Estimation

    Sep 14, 2026Yuxi Chen, Hamza Golubovic, Han Tong +2Influence FunctionInfluence

  27. MAxBench: A Multinomial Concept Recovery Benchmark

    Sep 14, 2026Divya Appapogu, Freya Behrens, Yonatan Belinkov +1InterpretabilitySteering

  28. Auditable Emergency Triage for Maternal and Newborn Care in India

    Sep 10, 2026Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini +7TriageInterpretability

  29. XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

    Sep 10, 2026Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein +1XaiExplainable Artificial Intelligence

  30. An Explainable Machine Learning Framework for Predicting Blood-Brain Barrier Permeability Using Molecular Descriptors

    Sep 9, 2026Fatemeh MahmoudiMolecular Property PredictionMoleculenet

  31. Do Reasoning Representations Help Humans Evaluate LLM Outputs?

    Sep 8, 2026Jaewoo Lim, Sungbok Shin, Sanghyun HongLLM Reasoning StrategiesLarge Reasoning Models

  32. LLM Layers Immediately Correct Each Other

    Sep 7, 2026Arjun Patrawala, Jiahai Feng, Erik Jones +1Transformer Residual StreamsTransformer Architectures

  33. When Superpixels Fail on Documents: A Study of Segmentation for LIME Explanations

    Sep 7, 2026Quentin Telnoff, Emanuela Boros, Mickaël Coustaty +3InterpretabilityDocument

  34. Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

    Sep 3, 2026Kevin Du, Alexander Hoyle, Laura Ruis +1Reasoning ChainReasoning Traces

  35. Witnesses Explain Anomalies

    Sep 3, 2026Lamine DiopGraph Anomaly DetectionInterpretability

  36. Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning

    Sep 1, 2026Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee +2Post-Disaster Damage AssessmentMultimodal Classification

  37. MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

    Aug 31, 2026Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang +5InterpretabilityMechanistic Interpretability

  38. Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

    Aug 31, 2026Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida +1Large Language Model CompressionTensor Networks

  39. Interpretable AI with Local Distillation

    Aug 24, 2026Erin Craig, Yiling Huang, Snigdha PanigrahiInterpretabilitySingular Learning Theory

  40. TabSOM: A tabular-to-image encoding method based on self-organizing maps

    Aug 13, 2026David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara +3Tabular LearningInterpretability

  41. Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

    Aug 13, 2026Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8DissociationLinguistics

  42. HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry

    Aug 12, 2026Haoran Pei, Zhao Su, Zetao Lin +6FuzzyHyperbolic Learning

  43. Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

    Aug 11, 2026Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo +3Drone DetectionUnmanned Aerial Vehicle Detection

  44. How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making

    Aug 10, 2026Adia Lumadjeng, Ilker Birbil, Erman AcarInterpretabilityFinancial Question Answering

  45. The Spectral Neuron

    Aug 8, 2026Alex ShtoffInterpretabilityAffine

  46. Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory

    Aug 7, 2026Kaustubh Kapil, Kishor P. UplaTransformer ArchitecturesVisual Geometry Grounded Transformer

  47. Faster Query-Key Learning Sharpens Attention in Self-Attention Models

    Aug 7, 2026Rahul Vashisht, Harish G. RamaswamyLinear AttentionInterpretability

  48. Evidential Rule Learning for Interpretable Classification with Abstention

    Aug 6, 2026Javier Fumanal-Idocin, Javier Andreu-PerezEvidential UncertaintyInterpretability

  49. Scaling Inherently Interpretable Language Models

    Aug 6, 2026Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7InterpretabilityDiffusion Language Models

  50. Benign interpolation and Occam's razor

    Aug 4, 2026Tom F. Sterkenburg, Daniel A. Herrmann, Jan-Willem RomeijnStochastic InterpolantsInterpretability

  51. Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

    Jul 31, 2026Tyler Ashoff, Jordan RoduArtificial Intelligence AlignmentMultimodal Alignment

  52. Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning

    Jul 29, 2026Alessio Cascione, Mattia Setzu, Cristiano Landi +2Decision TreesModel Selection