cs.LGMay 12, 2026

Do Fair Models Reason Fairly? Counterfactual Explanation Consistency for Procedural Fairness in Credit Decisions

Authors: Gideon PopoolaJohn Sheppard

Organizations: Gianforte School of Computing Montana State University Bozeman, MT, USA, 59715

Abstract

Machine learning algorithms in socially sensitive domains (e.g., credit decisions) often focus on equalizing predictive outcomes. However, satisfying these metrics does not guarantee that models use the same reasoning for different groups. We show that existing outcome-fair models can still apply fundamentally different reasoning to individuals, a hidden procedural bias'' missed by standard fairness metrics and algorithms. We propose Counterfactual Explanation Consistency (CEC), a framework that detects and mitigates this bias by aligning feature attributions between individuals and their counterfactual counterparts. Key contributions include a nearest-neighbor counterfactual generation method, a modified baseline for integrated gradient comparisons, an individual-level procedural fairness metric, and a corresponding training loss. We introduce a taxonomy identifying Regime B'' (same outcome, different reasoning) as a critical blind spot. Experiments on synthetic data, German Credit, Adult Income, and HMDA mortgage data demonstrate that outcome-fair baselines exhibit substantial hidden bias, while CEC substantially reduces it with modest utility cost.

Explore similar work

CardsList
  1. GESD: Beyond Outcome-Oriented Fairness

    May 14, 2026Gideon Popoola, John SheppardDisparitiesExplainability

  2. P2^2CE: Model-Agnostic Plausible Pareto-Optimal Counterfactual Explanations

    Jun 16, 2026Arthur Hendricks Mendes de Oliveira, Giovani Valdrighi, Marcos Medeiros RaimundoCounterfactual ExplanationOutliers