GFGE: Unifying Explainable AI Methods through an Interpretation Framework
Organizations: LIASD, Université Paris 8, 93200 Saint-Denis, France
Abstract
Explainable artificial intelligence (XAI) encompasses methods that draw on different sources of information and address different explanatory needs. A common framework is needed to describe how this information becomes evidence and is communicated as an explanation for a particular recipient. We propose the General Framework for Generating Explanations (GFGE), grounded in interpretative frameworks and the complementary activities of \emph{sense-reading} and \emph{sense-giving}. Its conceptual foundation is the Interpret/Explain Schema (IES), which connects an analyst's interpretation of system evidence, the communication of a selected account, and the recipient's interpretation of that account. GFGE operationalises this schema through five roles: data interpretation, model interpretation, output interpretation, optional post-hoc analysis, and aggregation. A role-typed operation graph records method-specific dependencies, while evidence records retain the sources, assumptions, and limitations of explanatory claims. The explanatory question, audience, and context guide the procedure. We instantiate GFGE for attribution, surrogate, counterfactual, concept and prototype, intrinsic rule, argumentation, and language-model methods. These instantiations show how intrinsic, post-hoc, and hybrid workflows can be represented through the same roles while preserving their distinct evidential requirements. GFGE provides a common basis for analysing explanation workflows, tracing communicated claims to their evidence, and identifying unresolved explanatory dependencies.
Figures & tables
| Review | Primary organising question | GFGE’s distinct analytical question |
|---|---|---|
| Schwalbe and Finzel [ 58 ] | Which properties distinguish methods and guide their selection? | Which interpretative operations and information flows generate a particular explanation? |
| Wang et al. [ 72 ] | Which stakeholders need which explanations at which lifecycle stage? | How does a selected method transform data, model, and output evidence into an account for that stakeholder? |
| Cousineau and Dara [ 17 ] ; Elsayed et al. [ 23 ] | How should techniques be classified or selected by mechanism and interpretability level? | How are intrinsic, post-hoc, and hybrid techniques instantiated in the same modular procedure? |
| Mangold et al. [ 43 ] | How should XAI systems be designed and evaluated with human users? | Which generation steps precede that evaluation, and where does recipient interpretation enter? |
| Cai et al. [ 12 ] | How can LLM interpretability support intervention and model improvement? | How do internal evidence, intervention tests, and communicated explanations fit into a framework spanning model families? |
| Role | Input output | Necessary assignment condition | Exclude and assign elsewhere |
|---|---|---|---|
| ID | Extracts feature meaning, measurement, domain constraints, or feasible changes. | Desired outcome belongs to ; a model response belongs to IO. | |
| IM | Directly reads a mechanism, parameter, rule, leaf, prototype, or circuit actually used by the predictor. | A fitted surrogate or attribution computation belongs to PH, even if model internals accelerate it. | |
| IO | Records observed predictions, scores, or queried contrasts produced by . | The requested alternative is in ; searching for inputs that attain it belongs to PH. | |
| PH | Runs an additional explanatory analysis: perturbation, attribution, surrogate fitting, intervention test, or candidate search. | Simply reading an existing model rule is IM; wording an existing claim is A. | |
| A | Selects, orders, and renders supported claims with provenance for audience . | Any new model query, statistical estimate, or factual inference must be recorded in IO, PH, ID, or IM first. |
| Family | Representative methods | GFGE roles and evidential issue |
|---|---|---|
| Feature attribution | SHAP, TreeSHAP, Integrated Gradients, DeepLIFT, LRP, Grad-CAM [ 41 , 40 , 63 , 60 , 7 , 59 ] | ID grounds features; IO records the prediction; PH computes contributions; A presents them. IM supplies structure for model-specific variants. |
| Local and global surrogates | LIME, surrogate trees, soft decision trees [ 55 , 18 , 27 ] | IO gathers model responses; PH fits an approximation; A reports its rules or coefficients with their scope. |
| Counterfactual and recourse | Wachter CF, DiCE, FACE, EECE, BRACE [ 69 , 49 , 52 , 73 , 25 ] | ID records mutable, plausible changes; fixes the alternative; IO checks outputs; PH searches candidates; A communicates the contrast. EECE also uses IM for forest structure. |
| Attribution-driven hybrid | SHAP-guided CF, CF-SHAP, recommendation CF [ 54 , 1 , 76 ] | PH composes attribution and search; ID constrains changes; IO records candidate model outputs; A combines findings and assumptions. |
| Concept and prototype | TCAV, ProtoPNet, ProtoTree, CBM-zero, FACE concept extraction [ 34 , 15 , 50 , 71 , 10 ] | ID grounds concepts or prototypes; IM exposes intrinsic concept mechanisms; PH tests post-hoc influence; A presents selected evidence. |
| Intrinsic rules and lists | Falling Rule Lists, CORELS, decision sets [ 70 , 3 , 36 ] | IM exposes predictive and active rules; IO records the result; ID grounds rule terms; A adapts the account to the question. |
| Method | Method-trait taxonomy | Stakeholder roadmap | GFGE trace and unresolved claim |
|---|---|---|---|
| CBM-zero [ 71 ] | Concept-based, model-access-dependent representation; local concept claims can be distinguished from global model traits. | Can pair a concept-level “what/why” question with an audience and stage. | PH constructs an invertible concept mapping; IM reads the converted predictor’s concept-to-output mechanism; IO records prediction preservation; ID must ground concept meaning. Recipient-specific A is unspecified. The order PH IM and the semantic gap are explicit. |
| ReCalX [ 19 ] | Perturbation-based, post-hoc analysis with reliability as a relevant method property. | Can identify a debugging need and method for a prediction question. | IO records scores under perturbation; PH checks calibration, adjusts the queried evidence, then recomputes an explanation. A downstream account and its audience are unspecified. An attribution without the calibration assumption remains a flagged claim. |
| FaithLM [ 16 ] | Language explanation with iterative intervention-based assessment. | Can connect a recipient’s question to a language explanation method. | IO records rationales produced by ; PH intervenes to test faithfulness and inform revision; A renders supported, recorded claims. This is a method-internal loop, not observed recipient feedback. Recipient understanding remains unspecified. |
| Workflow | Usual label | GFGE diagnosis |
|---|---|---|
| LIME/SHAP on latent features | Local, post-hoc | PH can compute scores while ID lacks a user-understandable mapping; A should state that limitation. |
| Attribution-guided counterfactual | Attribution or counterfactual | PH composes two operations; ID constrains changes, specifies the target, IO checks it, and A joins the claims. |
| EECE or causal recourse | Counterfactual, post-hoc | IM may supply forest structure; ID records feasible changes; PH searches candidates; A reports assumptions. |
| Intrinsic rule or argumentation model | Inherently interpretable | IM exposes reasoning structure, but IO selects the relevant contrast and A addresses the recipient. |
| Concept/prototype hybrid | Concept or prototype based | ID grounds semantics, IM or PH identifies influence, and A communicates the supported finding. |
| Generated LLM rationale | Natural-language explanation | New propositions are recorded as candidates and require evidence in ID, IM, IO, or PH; A renders supported claims. PH can test faithfulness to model behaviour. |