cs.CLOct 1, 2026

Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information

Authors: Tian Lan, Xiaoqing Cheng, Han Zhang, Jiang Li

Organizations: Kyoto University · Zhejiang University · Shanghai Jiao Tong University · Inner Mongolia University

Abstract

Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement

    Jan 29, 2026Jinhao Pan, Chahat Raj, Anjishnu Mukherjee +4Large Language Model BiasDebiasing

  2. BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection

    Aug 12, 2025Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3Large Language Model BiasBiases