cs.LGSep 11, 2026

Generative Interpretability via Scalable Neuro-Symbolic Models

Authors: Xiaocong Yang

Abstract

As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison to other interpretability research paradigms, and propose Neuro-Symbolic Models as a concrete instantiation.

Figures & tables

Explore similar work

CardsList
  1. Scaling Inherently Interpretable Language Models

    Aug 6, 2026Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7InterpretabilityDiffusion Language Models

  2. Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes

    Aug 3, 2026Maryam Rezaee, Pooriya Safaei, Maryam Asgarinezhad +1Post-HocFree-Energy Landscape